Every element, every page
Role, box, reading order, confidence and typed content: titles, text, tables, formulas, figures, captions, footnotes.
State-of-the-art AI parsing for .NET: PDFs, scans and images become typed, positioned elements, with charts read as data. Fully private.
{
"category": "figure",
"reading_index": 4,
"confidence": 0.97,
"bbox": [53, 70, 547, 329],
"content": {
"type": "chart",
"caption": "Revenue by region",
"text": "| Region | 2024 | 2025 |\n|---|---|---|\n| North | 12.4 | 15.1 |\n| South | 9.8 | 11.2 |"
}
}
Role, box, reading order, confidence and typed content: titles, text, tables, formulas, figures, captions, footnotes.
Flat extraction returns words. The answers your systems need live in the structure around them, and that is what the parser keeps.
Charts
Every bar, line and pie chart comes back as the table it plots, with its title, legend and panels.
Tables
Row and column spans survive as HTML, so a financial statement keeps its columns and its totals.
Order
Columns, sidebars, captions and footnotes read in the order a person reads them, never interleaved.
Position
Bounds and a confidence on each element make citations, highlighting and review possible.
Choose an effort level. The parser selects its models, verifies each one, and improves from one release to the next without a code change.
One parse renders every standard format, with figure pictures kept on request.
using LMKit.Document.Parsing; using var parser = new DocumentParser(ParsingEffort.Medium) { IncludeFigureImages = true }; ParsedDocument document = parser.Parse("annual-report.pdf"); File.WriteAllText("report.md", document.ToMarkdown()); File.WriteAllText("report.json", document.ToJson()); document.SaveHtml("report.html"); document.SaveDocLangArchive("report.dclx");
Every chart of a document, in reading order, as the data table it plots.
using LMKit.Document.Parsing; using var parser = new DocumentParser(ParsingEffort.High); ParsedDocument document = parser.Parse("filing.pdf"); foreach (ParsedPage page in document.Pages) foreach (ChartContent chart in page.Elements.Select(e => e.Content).OfType<ChartContent>()) { Console.WriteLine($"Page {page.PageNumber}: {chart.Caption}"); Console.WriteLine(chart.Text); // the plotted series as a table }
Store the lossless JSON and render any format later, without the source or the models.
using LMKit.Document.Parsing; // Parse once, store the result. File.WriteAllText("contract.json", parser.Parse("contract.pdf").ToJson()); // Later, on any machine: no parse, identical output. ParsedDocument stored = ParsedDocument.FromJson(File.ReadAllText("contract.json")); string html = stored.ToHtml();
Ship the model files with your application for air-gapped machines; each file is verified.
using LMKit.Document.Parsing; // On a connected machine: fetch what an effort level uses. foreach (var card in DocumentParser.GetModelCards(ParsingEffort.High)) card.Download(); // On the air-gapped machine: point the parser at the files. using var parser = new DocumentParser(ParsingEffort.High, modelFiles);
Every format renders from the same parse, so a heading, a table or a caption reads the same in all of them.
| Format | What it carries | Built for |
|---|---|---|
| Markdown | Headings, tables as HTML, formulas as LaTeX, charts as tables | LLM prompts and RAG chunks |
| JSON | Every element with category, box, reading order and confidence; a published schema | Pipelines and storage, in any language |
| HTML | Semantic markup with each element's category and box | Display, review and source highlighting |
| DocLang | The open AI-native document markup of the LF AI & Data Foundation | Exchange with DocLang tools and models |
LM-Kit One serves the same parser at POST /lmkit/v1/document-parsing,
on the command line and as an MCP tool for agents.
The fast universal converter across PDF, Office, email and HTML, with a vision model of your choice.
Index parsed pages with their tables and chart data, and answer with page citations.
Turn documents into typed fields with a confidence per field and a human-review flag.
Working console demos on GitHub, step-by-step how-to guides on the docs site, and the API reference for the classes used on this page.
Parse PDFs and scans into typed elements and export Markdown, JSON, HTML and DocLang.
Open on GitHub → DemoEvery chart of a folder of documents, read as data and written as CSV.
Open on GitHub → DemoQuestions over parsed documents, answered with page citations.
Open on GitHub → How-to guideEffort levels, elements, exports, figure pictures, bring-your-own models.
Read the guide → How-to guideCharts as data tables, panels and legends, exported as CSV.
Read the guide → API referenceAPI reference for the AI document parser.
Open the reference →