pdfextract API
    Preparing search index...

    Interface TextSpan

    A native text run or recognized OCR word with geometry and explicit provenance.

    interface TextSpan {
        bbox: Box;
        confidence?: number;
        fontName?: string;
        fontSize?: number;
        imageId?: string;
        quad?: Quad;
        source: "ocr" | "native";
        text: string;
    }
    Index
    bbox: Box

    Axis-aligned bounds in canonical page points (1/72 inch).

    confidence?: number

    OCR confidence in [0,1] when supplied; native text has no recognition confidence.

    fontName?: string

    Native PDF font identifier when available; not a claim that the font is installed.

    fontSize?: number

    Native text size in canonical page points when available.

    imageId?: string

    Source embedded appearance for regional OCR, when applicable.

    quad?: Quad

    Optional four-corner geometry in canonical page points.

    source: "ocr" | "native"

    Whether the span came from the PDF text layer or OCR.

    text: string

    Text content of the native run or OCR word.