Use case

Every number in the report. Even the ones in the charts.

Annual reports, filings and factsheets become structured data: tables with every span, charts as values, each figure traced to its page.

Tables with merged cells Charts read as data Page and box per figure Nothing leaves your network

Analysts re-key what the PDF already says.

Financial documents are built for print. Their numbers sit in spanning tables, in drawings and in footnotes, exactly where text extraction loses them.

Tables

Headers that span periods

Quarter and year-to-date columns share a header; flatten it and every value lands in the wrong column.

Charts

Numbers only drawn

Segment splits and trends often appear only as bars and lines, with no printed value to copy.

Context

Units and footnotes

"In thousands", negatives in parentheses and footnote markers change what a figure means.

From the PDF to your data model.

One engine reads the report, another fills your schema, and every value keeps the page and region it came from for audit.

01

Parse

Every page becomes typed elements: tables, charts, text, footnotes.

02

Tabulate

Tables keep their spans; charts come back as the table they plot.

03

Extract

Named figures fill your schema, each with a confidence.

04

Trace & load

Values load as JSON or CSV, with page and box for every number.

The chart is data. Read it that way.

No number is printed on these bars. The parser calibrates the axis, measures each bar and returns the table behind the picture, legend and units included.

  • Checked against the drawing. Values the page contradicts do not stand.
  • Labels from the print. Series and categories must be words the figure shows.
  • Every common form. Bars, stacks, lines, pies and multi-panel figures.

Every chart of a filing, as CSV.

The same parser runs inside .NET with LM-Kit.NET, or over REST, a command line and MCP with LM-Kit One.

LM-Kit.NET · FilingCharts.cs
using LMKit.Document.Parsing;

using var parser = new DocumentParser(ParsingEffort.High);
ParsedDocument filing = parser.Parse("annual-report.pdf");

foreach (ParsedPage page in filing.Pages)
foreach (ChartContent chart in page.Elements.Select(e => e.Content).OfType<ChartContent>())
{
    Console.WriteLine($"Page {page.PageNumber}: {chart.Caption}");
    Console.WriteLine(chart.Text);   // the plotted values as a table
}

Unpublished numbers stay unpublished.

Drafts, board packs and pre-release results are the documents that can least go to a cloud parser. Here, no page leaves your network.

Confidential

No third party sees a page.

Parsing runs on machines you administer, fully air-gapped if required. LM-Kit never receives the document.

Auditable

Every value has a source.

Each element keeps its page, box and confidence, so a reviewer jumps straight to the region behind a figure.

Predictable

No per-page API charge.

Backfill ten years of filings on hardware you already run; reporting season is a scheduling question.

Frequently asked questions.

Can it read values that only appear in a chart?

Yes. The parser returns every bar, line, pie and stacked chart as the data table it plots, with its legend and units. Values are measured from the drawing and checked against it, so a figure printed nowhere as text still becomes a number.

Does it keep multi-level table headers?

Yes. Row and column spans survive as HTML in the Markdown and JSON output, so a header that covers two periods stays above both columns, and negatives in parentheses and bold totals are kept.

How do we trace a number back to the report?

Every element carries its page, bounding box, reading position and confidence. A reviewer or a downstream system can open the exact region a value came from.

Do the documents leave our network?

No. The parser runs inside your .NET application with LM-Kit.NET, or on your servers with LM-Kit One, including fully air-gapped machines. LM-Kit never receives the documents.

Financial reports

Parse your hardest report on your own hardware.