Use case

Papers your LLM can read. Formulas, figures and all.

Two-column papers, equations and plots become structured Markdown and JSON: formulas as LaTeX, plots as data, every claim traced to its page.

Columns in reading order Formulas as LaTeX Plots read as data Cited to page and box

A paper is not a wall of text.

Literature search and research assistants fail on the same pages: two columns, equations and figures that text extraction turns into noise.

Columns

Two arguments, interleaved

Read line by line across the page, two columns merge into sentences no author wrote.

Formulas

Equations as symbol soup

Subscripts, Greek letters and fractions come out flattened, and the result loses its meaning.

Figures

Results trapped in plots

The key result is often a curve or a bar chart, invisible to anything that reads only text.

The structure of the paper, kept.

The parser returns the paper as typed elements in reading order, ready for Markdown, JSON, HTML or DocLang.

Order

Read like a person

Columns, captions and sidebars in the order the author meant, footnotes and running heads kept apart.

LaTeX

Formulas you can render

Display equations come back as LaTeX, ready for a renderer, a notebook or an LLM prompt.

Data

Plots as tables

Bar charts, line plots and multi-panel figures return the values they plot, with legend and units.

Tables

Results tables, whole

Grouped headers and merged cells survive as HTML spans, so each value stays in its column.

Outline

Sections across pages

Heading levels are ranked across the whole paper, so a section keeps its level on every page.

Grounding

Every claim has a source

Each element carries its page and box, so an answer can cite the exact paragraph, table or figure.

Ask the literature. Get cited answers.

Parse once, index the elements, and a research assistant answers from the papers themselves, pointing to the page and region behind each claim.

  • Chunks that follow the paper. Split on sections; a table or a figure stays whole.
  • Values you can retrieve. Plot data and table cells become text an index matches.
  • Answers that cite. Page and box travel with every chunk into the answer.
  • Built in. Document RAG runs on the same engine.

A folder of papers, LLM-ready.

The same parser runs inside .NET with LM-Kit.NET, or over REST, a command line and MCP with LM-Kit One.

LM-Kit One · terminal
# Every paper in ./papers as Markdown: LaTeX formulas, plots as tables, sections kept
lmkit run document-parsing --in ./papers --out ./markdown --effort High --output-format Markdown

Frequently asked questions.

How are equations returned?

Display formulas come back as LaTeX in the Markdown, JSON, HTML and DocLang output, so they can be rendered, searched or passed to an LLM without losing their structure.

Can it read the values in a plot?

Yes. Bar charts, line plots, pies and multi-panel figures are returned as the data table they plot, with legend and units, so a result shown only as a curve becomes values an index can find.

Does it handle two-column layouts?

Yes. Elements come back in the order a person reads them, with captions bound to their figures and footnotes and running heads kept apart from the body.

Can answers cite the paper precisely?

Every element carries its page, bounding box and confidence. Parsed papers feed Document RAG, which answers with the page and region behind each claim.

Scientific papers

Make your literature readable by AI, privately.