Columns
Two arguments, interleaved
Read line by line across the page, two columns merge into sentences no author wrote.
Two-column papers, equations and plots become structured Markdown and JSON: formulas as LaTeX, plots as data, every claim traced to its page.
Literature search and research assistants fail on the same pages: two columns, equations and figures that text extraction turns into noise.
Columns
Read line by line across the page, two columns merge into sentences no author wrote.
Formulas
Subscripts, Greek letters and fractions come out flattened, and the result loses its meaning.
Figures
The key result is often a curve or a bar chart, invisible to anything that reads only text.
The parser returns the paper as typed elements in reading order, ready for Markdown, JSON, HTML or DocLang.
Order
Columns, captions and sidebars in the order the author meant, footnotes and running heads kept apart.
LaTeX
Display equations come back as LaTeX, ready for a renderer, a notebook or an LLM prompt.
Data
Bar charts, line plots and multi-panel figures return the values they plot, with legend and units.
Tables
Grouped headers and merged cells survive as HTML spans, so each value stays in its column.
Outline
Heading levels are ranked across the whole paper, so a section keeps its level on every page.
Grounding
Each element carries its page and box, so an answer can cite the exact paragraph, table or figure.
Parse once, index the elements, and a research assistant answers from the papers themselves, pointing to the page and region behind each claim.
The same parser runs inside .NET with LM-Kit.NET, or over REST, a command line and MCP with LM-Kit One.
# Every paper in ./papers as Markdown: LaTeX formulas, plots as tables, sections kept
lmkit run document-parsing --in ./papers --out ./markdown --effort High --output-format Markdown
using LMKit.Document.Parsing; using var parser = new DocumentParser(ParsingEffort.High); ParsedDocument paper = parser.Parse("paper.pdf"); File.WriteAllText("paper.md", paper.ToMarkdown()); // for the LLM File.WriteAllText("paper.json", paper.ToJson()); // pages, boxes, confidence
Preprints, lab notebooks and patent drafts are intellectual property. The parser runs on your own machines, air-gapped if needed.
Edge
Embed the parser in a .NET tool on the machine where the papers already are.
Edge deploymentLocal
Serve it from LM-Kit One to any language, notebook or agent on your network.
Deploy LM-Kit OneSovereign
No page leaves your infrastructure, and LM-Kit never receives a document.
SovereigntyDisplay formulas come back as LaTeX in the Markdown, JSON, HTML and DocLang output, so they can be rendered, searched or passed to an LLM without losing their structure.
Yes. Bar charts, line plots, pies and multi-panel figures are returned as the data table they plot, with legend and units, so a result shown only as a curve becomes values an index can find.
Yes. Elements come back in the order a person reads them, with captions bound to their figures and footnotes and running heads kept apart from the body.
Every element carries its page, bounding box and confidence. Parsed papers feed Document RAG, which answers with the page and region behind each claim.
Scientific papers