Solutions · Document Intelligence · Document parsing

Parse any document. Keep every structure.

State-of-the-art AI parsing for .NET: PDFs, scans and images become typed, positioned elements, with charts read as data. Fully private.

Charts read as data Tables with merged cells JSON, HTML, DocLang 100% on-device
report.json
{
  "category": "figure",
  "reading_index": 4,
  "confidence": 0.97,
  "bbox": [53, 70, 547, 329],
  "content": {
    "type": "chart",
    "caption": "Revenue by region",
    "text": "| Region | 2024 | 2025 |\n|---|---|---|\n| North | 12.4 | 15.1 |\n| South | 9.8 | 11.2 |"
  }
}

Every element, every page

Role, box, reading order, confidence and typed content: titles, text, tables, formulas, figures, captions, footnotes.

A text layer is not a document.

Flat extraction returns words. The answers your systems need live in the structure around them, and that is what the parser keeps.

Charts

Read as data

Every bar, line and pie chart comes back as the table it plots, with its title, legend and panels.

Tables

Merged cells kept

Row and column spans survive as HTML, so a financial statement keeps its columns and its totals.

Order

Reading order

Columns, sidebars, captions and footnotes read in the order a person reads them, never interleaved.

Position

A box for every element

Bounds and a confidence on each element make citations, highlighting and review possible.

One setting. The engine does the rest.

Choose an effort level. The parser selects its models, verifies each one, and improves from one release to the next without a code change.

  • Low. The least time per page, for high-volume ingestion where throughput leads.
  • Medium. The balance of time and fidelity that suits most documents; the default.
  • High. The most faithful read, for documents where every table, chart and heading counts.
  • Verified models. Only the models the parser is tested with run; any other is refused before a page is read.
  • Offline ready. Download the models once, or ship the files with your application.

A few lines from file to data.

One parse renders every standard format, with figure pictures kept on request.

ParseAndExport.cs
using LMKit.Document.Parsing;

using var parser = new DocumentParser(ParsingEffort.Medium)
{
    IncludeFigureImages = true
};

ParsedDocument document = parser.Parse("annual-report.pdf");

File.WriteAllText("report.md",   document.ToMarkdown());
File.WriteAllText("report.json", document.ToJson());
document.SaveHtml("report.html");
document.SaveDocLangArchive("report.dclx");

Standard formats, not a proprietary result.

Every format renders from the same parse, so a heading, a table or a caption reads the same in all of them.

Format What it carries Built for
Markdown Headings, tables as HTML, formulas as LaTeX, charts as tables LLM prompts and RAG chunks
JSON Every element with category, box, reading order and confidence; a published schema Pipelines and storage, in any language
HTML Semantic markup with each element's category and box Display, review and source highlighting
DocLang The open AI-native document markup of the LF AI & Data Foundation Exchange with DocLang tools and models

Embed it, or serve it.

LM-Kit One serves the same parser at POST /lmkit/v1/document-parsing, on the command line and as an MCP tool for agents.

LM-Kit One, the Private AI Application Server

Parsing feeds everything else.

Document to Markdown

The fast universal converter across PDF, Office, email and HTML, with a vision model of your choice.

Markdown conversion

Document RAG

Index parsed pages with their tables and chart data, and answer with page citations.

Document RAG

Structured extraction

Turn documents into typed fields with a confidence per field and a human-review flag.

Structured extraction

Every structure kept. No cloud.

Free Download Read the guide