pdfextract API
    Preparing search index...

    Interface TesseractOcrOptions

    Worker-pool and asset configuration. No worker or model is loaded until recognition begins.

    interface TesseractOcrOptions {
        assets?: {
            coreBaseUrl?: string | URL;
            languageDataBaseUrl?: string | URL;
            workerUrl?: string | URL;
        };
        concurrency?: number;
        languages: readonly string[];
    }
    Index
    assets?: {
        coreBaseUrl?: string | URL;
        languageDataBaseUrl?: string | URL;
        workerUrl?: string | URL;
    }

    Self-hosted asset locations; languageDataBaseUrl is required when recognition starts.

    Type Declaration

    • OptionalcoreBaseUrl?: string | URL

      Directory containing the shipped Tesseract LSTM core JS/WASM variants; defaults to the packaged assets/core. Supports Node filesystem paths and file URLs.

    • OptionallanguageDataBaseUrl?: string | URL

      Directory containing language.traineddata.gz files. No implicit language-model CDN is used.

    • OptionalworkerUrl?: string | URL

      Worker script location. Defaults to the packaged worker next to this module; bundled browser apps should supply the hosted worker.min.js URL.

    concurrency?: number

    Maximum reusable worker count, 1–8; defaults to one.

    languages: readonly string[]

    Nonempty Tesseract language identifiers, for example ['eng', 'deu']. Supply matching traineddata.gz files.