{"id":2137,"date":"2026-09-29T08:30:39","date_gmt":"2026-09-29T08:30:39","guid":{"rendered":"https:\/\/cms.research.wpp.com\/?post_type=research_feed&#038;p=2137"},"modified":"2026-09-29T09:52:52","modified_gmt":"2026-09-29T09:52:52","slug":"mineru-the-hard-part-isnt-the-parsing","status":"publish","type":"research_feed","link":"https:\/\/cms.research.wpp.com\/?research_feed=mineru-the-hard-part-isnt-the-parsing","title":{"rendered":"Evaluating MinerU for PDF-to-Markdown Extraction"},"content":{"rendered":"\n<div class=\"wp-block-group has-global-padding is-layout-constrained wp-block-group-is-layout-constrained\">\n<div class=\"wp-block-group has-global-padding is-layout-constrained wp-block-group-is-layout-constrained\">\n<p class=\"wp-block-paragraph\">MinerU is an open-source document parser from OpenDataLab, the open-data group at Shanghai Artificial Intelligence Laboratory [1]. It converts PDFs, images and Microsoft Office files into Markdown, a plain-text format with simple marks for headings, lists and tables. It can also produce JSON that records the page layout.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">We installed MinerU 4.0 on a laptop and ran it on five documents: an arXiv preprint in mathematics, a SIGMOD paper on data management, MinerU&#8217;s own paper, a photographed page and a typeset dictionary sample.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Our results show that MinerU produces results that are clearly superior to simple PDF-parsing baselines, respecting much more of the underlying content and structure. This makes it a useful module for any application that relies on the accurate comprehension of unstructured data .<\/p>\n<\/div>\n<\/div>\n\n\n\n<hr class=\"wp-block-separator has-alpha-channel-opacity\"\/>\n\n\n\n<h2 class=\"wp-block-heading\">Why not just extract the text?<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">We use one page throughout the article to make the problem concrete: the first page of the dictionary sample, which we call the dictionary page (Figure 1). At the top, a blue bar holds the running head, a line repeated on every page. The same phrase appears underneath as the document title. The main part of the page is a list of dictionary entries in two columns. At the bottom are a page number and a short block of CSS, the code used to style web pages. From this page we want the entries in reading order, the CSS kept as code, and no running head or page number in the text.<\/p>\n\n\n\n<figure class=\"wp-block-image aligncenter size-large is-resized\"><img loading=\"lazy\" decoding=\"async\" width=\"725\" height=\"1024\" src=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/dictionary-page-725x1024.png\" alt=\"\" class=\"wp-image-2146\" style=\"aspect-ratio:0.70801317233809;width:479px;height:auto\" srcset=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/dictionary-page-725x1024.png 725w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/dictionary-page-212x300.png 212w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/dictionary-page-768x1085.png 768w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/dictionary-page.png 1000w\" sizes=\"auto, (max-width: 725px) 100vw, 725px\" \/><figcaption class=\"wp-element-caption\">The first page of the dictionary sample.<\/figcaption><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">Our second example is the photographed page, a printed page from Google&#8217;s ML Kit documentation that was photographed with a phone and saved as a PDF (Figure 2).<\/p>\n\n\n\n<figure class=\"wp-block-image aligncenter size-large is-resized\"><img loading=\"lazy\" decoding=\"async\" width=\"791\" height=\"1024\" src=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/photographed-page-1-791x1024.png\" alt=\"\" class=\"wp-image-2153\" style=\"aspect-ratio:0.7724695447145343;width:470px;height:auto\" srcset=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/photographed-page-1-791x1024.png 791w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/photographed-page-1-232x300.png 232w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/photographed-page-1-768x994.png 768w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/photographed-page-1-1187x1536.png 1187w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/photographed-page-1.png 1190w\" sizes=\"auto, (max-width: 791px) 100vw, 791px\" \/><figcaption class=\"wp-element-caption\">The photographed page: Google ML Kit documentation, taken with a phone.<\/figcaption><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">As a baseline, we ran both pages through <code>pypdfium2<\/code>, a Python library that extracts the text stored in a PDF. A PDF usually stores a page as many small pieces of text, each with its own position. <code>pypdfium2<\/code> returns them in the order the file lists them, which is not always the order a person reads them.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">On the photographed page, <code>pypdfium2<\/code> returned zero characters. That is expected because the PDF contains only an image and no text layer for an extractor to read. It returned an empty string with no warning and no error.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">On the dictionary page, it returned 2,707 characters, but in the wrong order. It listed 29 headwords, the words being defined, from <code>Peace, n.<\/code> to <code>Pedantic, a.<\/code>, before giving the first definition. Its first lines were:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code>Multi-column Sample\nPeace, n.\nPeaceable, a.\nPeaceful, a.\nPeace-maker, n.\n   ... 24 more headwords ...\nPedantic, a.\nMulti-column Sample\n1. Dictionary Layout\nThe following example shows how to create a basic multi-column document\n1. Calm, repose, quiet, tran?quillity, stillness, silence.\n<\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\">The output has three problems:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>The definition of <code>Peace<\/code> appears only after all 29 headwords.<\/li>\n\n\n\n<li><code>Multi-column Sample<\/code> appears twice, because the extractor cannot tell the running head from the title.<\/li>\n\n\n\n<li><code>tranquillity<\/code> contains a <code>?<\/code>, our mark for a character the extractor could not read, at the point where the word breaks across two lines in Figure 1. MinerU returned the word correctly.<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">Apart from that character, all the text is present. The problem is the structure: which column a line belongs to, and which repeated line is a running head rather than a title. MinerU is designed to recover that structure, along with tables, equations and captions [1]. We tested it on the five documents listed in Table 1.<\/p>\n\n\n\n<hr class=\"wp-block-separator has-alpha-channel-opacity\"\/>\n\n\n\n<h2 class=\"wp-block-heading\">What is MinerU?<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Instead of reading the text stored in the PDF, MinerU works from what the page looks like. It turns each page into an image, finds the parts of the page, and sends each part to a model built for that kind of content. A model here is a neural network trained for a particular task. An equation, for example, goes to a model that writes it out as LaTeX.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The first step, layout detection, divides the page into regions and labels each one. For the dictionary page, it produced these labels:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code>header           \"Multi-column Sample\"\nparagraph_title  \"Multi-column Sample\"\nparagraph_title  \"1. Dictionary Layout\"\ntext             \"The following example shows how to create a basic...\"\nref_text         \"Peace, n.1. Calm, repose, quiet, tranquillity...\"   (x29 entries)\ntext             \"Excerpt from \"Dictionary of English Synonymes\" by...\"\nfooter           \"div.dictionary {height 6in; \/* if not set the box...\"\npage_number      \"1\"\nfooter           \"www.pdfreactor.com\/manual\"\n<\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\">The label determines how each region is processed next:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li><code>text<\/code>, <code>ref_text<\/code> and <code>paragraph_title<\/code> regions go through optical character recognition (OCR), which reads characters from the image rather than from the PDF&#8217;s stored text. OCR is how MinerU reads the photographed page.<\/li>\n\n\n\n<li><code>equation<\/code> regions go to a formula model, which writes LaTeX.<\/li>\n\n\n\n<li><code>table<\/code> regions go to a table model. It writes HTML, the markup language used for web pages, when cells span several rows or columns, and a Markdown table otherwise.<\/li>\n\n\n\n<li><code>image<\/code> and <code>chart<\/code> regions are cut out of the page and kept as pictures.<\/li>\n\n\n\n<li><code>header<\/code>, <code>footer<\/code> and <code>page_number<\/code> regions are discarded.<\/li>\n<\/ul>\n\n\n\n<p class=\"has-text-align-left wp-block-paragraph\">Only two of these apply to the dictionary page: OCR and discarding. Regions are discarded based on their labels. That keeps running heads and page numbers out of the text, but it can also remove content that was given the wrong label. Two regions on the dictionary page are labelled <code>footer<\/code>: the web address at the bottom, which is a real footer, and the CSS listing, which is not. MinerU discarded both.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">This is the first part of MinerU&#8217;s Markdown output for the dictionary page:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code>## Multi-column Sample\n\n## 1. Dictionary Layout\n\nThe following example shows how to create a basic multi-column document\nwith two columns, balanced content and a small gap.\n\nPeace, n.1. Calm, repose, quiet, tranquillity, stillness, silence. 2.\nAmity, concord, harmony, truce, ARMISTICE.\n<\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\">Unlike the <code>pypdfium2<\/code> output, this one puts the first definition straight after its headword, keeps <code>tranquillity<\/code> intact, and prints the title only once. The running head was labelled <code>header<\/code> and discarded. The title was labelled <code>paragraph_title<\/code> and kept.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">How much of this processing runs depends on a setting called the tier. We used the default for PDFs, <code>standard<\/code>, which adds a vision-language model, a larger model that reads image and text together, for regions the single-task models cannot resolve.<\/p>\n\n\n\n<hr class=\"wp-block-separator has-alpha-channel-opacity\"\/>\n\n\n\n<h2 class=\"wp-block-heading\">How we tested it<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">We chose five documents, each testing a different part of PDF parsing (Table 1). The arXiv preprint [2] is a mathematics paper with many fraktur letters. Fraktur is a Gothic-style alphabet that mathematicians use for some symbols. The SIGMOD paper [3] is from the ACM SIGMOD conference on data management.<\/p>\n\n\n\n<div class=\"compact-table pdi-table\" style=\"overflow-x:auto; margin-bottom:0 !important;\">\n  <table style=\"width:100%; table-layout:auto; border-collapse:collapse; font-size:16px; line-height:1.35;\">\n    <caption style=\"caption-side: bottom; text-align: left; margin-top: 10px; line-height: 1.4; font-size: 14px; color: #555;\">\n      <strong>Table 1.<\/strong> The five documents we tested.\n    <\/caption>\n    <colgroup>\n      <col style=\"width:30%;\">\n      <col style=\"width:20%;\">\n      <col style=\"width:50%;\">\n    <\/colgroup>\n    <thead>\n      <tr>\n        <th style=\"border: 1px solid #ccc; padding: 10px; text-align: left;\"><strong>Document<\/strong><\/th>\n        <th style=\"border: 1px solid #ccc; padding: 10px; text-align: left;\"><strong>Pages<\/strong><\/th>\n        <th style=\"border: 1px solid #ccc; padding: 10px; text-align: left;\"><strong>What it tests<\/strong><\/th>\n      <\/tr>\n    <\/thead>\n    <tbody>\n      <tr>\n        <td style=\"border: 1px solid #ccc; padding: 10px;\">Dictionary sample<\/td>\n        <td style=\"border: 1px solid #ccc; padding: 10px;\">6<\/td>\n        <td style=\"border: 1px solid #ccc; padding: 10px;\">Reading order across two columns<\/td>\n      <\/tr>\n      <tr>\n        <td style=\"border: 1px solid #ccc; padding: 10px;\">arXiv preprint [2]<\/td>\n        <td style=\"border: 1px solid #ccc; padding: 10px;\">61<\/td>\n        <td style=\"border: 1px solid #ccc; padding: 10px;\">Dense mathematics and fraktur letters<\/td>\n      <\/tr>\n      <tr>\n        <td style=\"border: 1px solid #ccc; padding: 10px;\">SIGMOD paper [3]<\/td>\n        <td style=\"border: 1px solid #ccc; padding: 10px;\">25<\/td>\n        <td style=\"border: 1px solid #ccc; padding: 10px;\">Figures and captions<\/td>\n      <\/tr>\n      <tr>\n        <td style=\"border: 1px solid #ccc; padding: 10px;\">Photographed page<\/td>\n        <td style=\"border: 1px solid #ccc; padding: 10px;\">1<\/td>\n        <td style=\"border: 1px solid #ccc; padding: 10px;\">OCR on a page with no text layer<\/td>\n      <\/tr>\n      <tr>\n        <td style=\"border: 1px solid #ccc; padding: 10px;\">MinerU&#8217;s own paper [4]<\/td>\n        <td style=\"border: 1px solid #ccc; padding: 10px;\">57<\/td>\n        <td style=\"border: 1px solid #ccc; padding: 10px;\">A long paper with many figures and tables<\/td>\n      <\/tr>\n      <tr>\n        <td style=\"border: 1px solid #ccc; padding: 10px;\">Total<\/td>\n        <td style=\"border: 1px solid #ccc; padding: 10px;\">150<\/td>\n        <td style=\"border: 1px solid #ccc; padding: 10px;\"><\/td>\n      <\/tr>\n    <\/tbody>\n  <\/table>\n<\/div>\n\n\n\n<p class=\"wp-block-paragraph\">Table 2 lists the machine and software we used..<\/p>\n\n\n\n<div class=\"compact-table pdi-table\" style=\"overflow-x:auto; margin-bottom:0 !important;\">\n  <table style=\"width:100%; table-layout:auto; border-collapse:collapse; font-size:16px; line-height:1.35;\">\n    <caption style=\"caption-side: bottom; text-align: left; margin-top: 10px; line-height: 1.4; font-size: 14px; color: #555;\">\n      <strong>Table 2.<\/strong> Experimental setup.\n    <\/caption>\n    <colgroup>\n      <col style=\"width:25%;\">\n      <col style=\"width:75%;\">\n    <\/colgroup>\n    <thead>\n      <tr>\n        <th style=\"border: 1px solid #ccc; padding: 10px; text-align: left;\"><strong>Item<\/strong><\/th>\n        <th style=\"border: 1px solid #ccc; padding: 10px; text-align: left;\"><strong>Value<\/strong><\/th>\n      <\/tr>\n    <\/thead>\n    <tbody>\n      <tr>\n        <td style=\"border: 1px solid #ccc; padding: 10px;\">Machine<\/td>\n        <td style=\"border: 1px solid #ccc; padding: 10px;\">Apple M5 Pro, 48 GB unified memory, <code>arm64<\/code><\/td>\n      <\/tr>\n      <tr>\n        <td style=\"border: 1px solid #ccc; padding: 10px;\">Acceleration<\/td>\n        <td style=\"border: 1px solid #ccc; padding: 10px;\">GPU via Apple Metal (<code>device=mps<\/code>) for the small models, <code>llama.cpp<\/code> for the vision-language model<\/td>\n      <\/tr>\n      <tr>\n        <td style=\"border: 1px solid #ccc; padding: 10px;\">MinerU<\/td>\n        <td style=\"border: 1px solid #ccc; padding: 10px;\"><code>4.0.0<\/code>, reported by <code>mineru version --json<\/code><\/td>\n      <\/tr>\n      <tr>\n        <td style=\"border: 1px solid #ccc; padding: 10px;\">Python<\/td>\n        <td style=\"border: 1px solid #ccc; padding: 10px;\">3.13.15, in a <code>uv<\/code> virtual environment<\/td>\n      <\/tr>\n      <tr>\n        <td style=\"border: 1px solid #ccc; padding: 10px;\">Tier<\/td>\n        <td style=\"border: 1px solid #ccc; padding: 10px;\"><code>standard<\/code>: the single-task models plus a 1.2-billion-parameter vision-language model<\/td>\n      <\/tr>\n      <tr>\n        <td style=\"border: 1px solid #ccc; padding: 10px;\">Date<\/td>\n        <td style=\"border: 1px solid #ccc; padding: 10px;\">17 September 2026<\/td>\n      <\/tr>\n    <\/tbody>\n  <\/table>\n<\/div>\n\n\n\n<p class=\"wp-block-paragraph\">We used the default settings, with no configuration file and no tuning.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">By default MinerU writes each figure into the Markdown file as base64, a way of encoding an image as text. MinerU&#8217;s own paper produced a single 17 MB Markdown file that way. With <code>--format zip<\/code>, MinerU saves the 117 figures as separate JPEG files and the Markdown file drops to 193 KB.<\/p>\n\n\n\n<hr class=\"wp-block-separator has-alpha-channel-opacity\"\/>\n\n\n\n<h2 class=\"wp-block-heading\">What it got right<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>MinerU kept the dictionary sample&#8217;s two columns in reading order.<\/strong> In Figure 1, <code>Pearlash<\/code> ends the left column and <code>Pearl-white<\/code> starts the right one. MinerU&#8217;s output keeps them in that order, each with its own definition:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code>Pearlash, n. Sub-carbonate of potassa (impure), calcined potash. PEARL-WHITE.\n\nPearl-white, n.Pearl-powder, submuriate of bismuth.\n<\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Running heads, footers and page numbers stayed out of the dictionary sample&#8217;s Markdown.<\/strong> Across the six pages, MinerU labelled 6 headers, 7 footers and 6 page numbers and discarded them all, so neither the web address nor the page numbers appear in the output. One of the footers was the CSS listing described in &#8220;Where it breaks&#8221;.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>MinerU read every word on the photographed page correctly.<\/strong> It also marked the two blue headings in Figure 2, <code>ML Kit<\/code> and <code>Document scanner<\/code>, as Markdown headings.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>MinerU correctly rebuilt a table with a two-level header and sideways row labels.<\/strong> The table comes from MinerU&#8217;s own paper and is shown in Figure 3. In its header, <code>Stage-0<\/code> spans two sub-columns, <code>a<\/code> and <code>b<\/code>. Down the left side, four group labels, <code>Vision<\/code>, <code>Data<\/code>, <code>Model<\/code> and <code>Training<\/code>, are printed sideways.<\/p>\n\n\n\n<figure class=\"wp-block-image size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"964\" height=\"431\" src=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/rotated-table.png\" alt=\"\" class=\"wp-image-2149\" srcset=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/rotated-table.png 964w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/rotated-table-300x134.png 300w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/rotated-table-767x343.png 767w\" sizes=\"auto, (max-width: 964px) 100vw, 964px\" \/><figcaption class=\"wp-element-caption\">The two-level header table in MinerU&#8217;s own paper, as printed [4].<\/figcaption><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">MinerU returned the table as HTML. The start of its output, reformatted for reading, is:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code>&lt;table&gt;\n  &lt;tr&gt;\n    &lt;td rowspan=\"2\" colspan=\"2\"&gt;&lt;\/td&gt;\n    &lt;td colspan=\"2\"&gt;Stage-0&lt;\/td&gt;\n    &lt;td rowspan=\"2\"&gt;Stage-1&lt;\/td&gt;\n    &lt;td rowspan=\"2\"&gt;Stage-2&lt;\/td&gt;\n  &lt;\/tr&gt;\n  &lt;tr&gt;\n    &lt;td&gt;a&lt;\/td&gt;\n    &lt;td&gt;b&lt;\/td&gt;\n  &lt;\/tr&gt;\n  &lt;tr&gt;\n    &lt;td rowspan=\"2\"&gt;Vision&lt;\/td&gt;\n    &lt;td&gt;Max Resolution&lt;\/td&gt;\n    &lt;td&gt;$2048 \\times 28 \\times 28$&lt;\/td&gt;\n    ...\n<\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\">We checked the output cell by cell against the printed table, and every cell was correct, including the numbers written in LaTeX. Each sideways label became a group spanning the right rows: 2 for <code>Vision<\/code>, 2 for <code>Data<\/code>, 3 for <code>Model<\/code> and 4 for <code>Training<\/code>.<\/p>\n\n\n\n<p class=\"is-style-default wp-block-paragraph\"><strong>MinerU recovered the structure of 9 tables that appear in its own paper [4] only as images.<\/strong> MinerU returned 17 tables for its own paper, but only 8 are typeset tables in the PDF. The other 9 are inside images that the paper shows as examples of its model&#8217;s output. All 9 came back with the right rows and columns.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>All 192 display equations in the arXiv preprint [2] came back as well-formed LaTeX.<\/strong> Display equations are the ones printed on a line of their own. A script checked every one for structural errors, such as an unclosed brace or a <code>\\left<\/code> without its <code>\\right<\/code>. That check cannot show whether an equation matches the printed one, so we also compared equations by hand. Figure 4 shows one of them, from page 8.<\/p>\n\n\n\n<figure class=\"wp-block-image size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"1000\" height=\"109\" src=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/equation-source.png\" alt=\"\" class=\"wp-image-2150\" srcset=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/equation-source.png 1000w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/equation-source-761x83.png 761w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/equation-source-300x33.png 300w\" sizes=\"auto, (max-width: 1000px) 100vw, 1000px\" \/><figcaption class=\"wp-element-caption\">A display equation on page 8 of the arXiv preprint, as printed [2].<\/figcaption><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">And here is what MinerU returned for it:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code>$$\\mathfrak {g} = \\mathfrak {k} \\oplus \\mathfrak {p}.$$\n<\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\">All three fraktur letters are correct, and the circled plus, <code>\u2295<\/code>, became <code>\\oplus<\/code>, its LaTeX command, rather than a plain plus sign.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Figure captions, footnotes and author affiliations stayed with the content they belong to.<\/strong> All the captions in the SIGMOD paper came out next to their corresponding figures, in the original order. Footnotes were labelled as footnotes rather than mixed into the main text, and the sixty-author list in MinerU&#8217;s own paper kept its superscript affiliations and links.<\/p>\n\n\n\n<hr class=\"wp-block-separator has-alpha-channel-opacity\"\/>\n\n\n\n<h2 class=\"wp-block-heading\">Where it breaks<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>First, MinerU deleted one CSS listing from the dictionary sample and left no sign of the gap.<\/strong> The dictionary sample prints five CSS listings. MinerU labelled four as code and kept them. The fifth, the <code>div.dictionary<\/code> listing at the bottom of the dictionary page (Figure 1), was labelled a footer and discarded, with nothing in its place. This is the only case of lost content we found.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>In addition, MinerU turned some ordinary text in the arXiv preprint into formulas.<\/strong> In LaTeX, dollar signs mark where a formula starts and ends. Citations such as <code>[Ada+20]<\/code> came back as <code>$\\left[ \\mathrm{Ada} + 20 \\right]$<\/code>, which is still readable but is now a formula rather than a citation. MinerU occasionally moved punctuation inside formulas, as in <code>In Section $5 ,$ we formulate<\/code> and <code>Using $\\beta ,$ we define<\/code>.<\/p>\n\n\n\n<hr class=\"wp-block-separator has-alpha-channel-opacity\"\/>\n\n\n\n<h2 class=\"wp-block-heading\">Conclusion<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">MinerU handles document layouts that caused the text-only <code>pypdfium2<\/code> baseline to fail. It reads photographed pages with no text layer, preserves natural reading order across columns, reconstructs tables with multi-level and sideways headers, and converts display equations into well-formed LaTeX.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">In some cases, its processing can change or remove content. We found one CSS listing that was deleted without a placeholder, and ordinary text and citations that were formatted as mathematical formulas.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Overall, MinerU was shown to be clearly superior to simple PDF parsing, making it a useful module for any application that requires high-fidelity representation of unstructured textual data.<\/p>\n\n\n\n<hr class=\"wp-block-separator has-alpha-channel-opacity\"\/>\n\n\n\n<h2 class=\"wp-block-heading\">References<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">[1] Wang et al. <em>MinerU: An Open-Source Solution for Precise Document Content Extraction<\/em>. <a href=\"https:\/\/arxiv.org\/abs\/2409.18839\">https:\/\/arxiv.org\/abs\/2409.18839<\/a><\/p>\n\n\n\n<p class=\"wp-block-paragraph\">[2] Afentoulidis et al. <em>Dirac operators for algebraic families<\/em>. <a href=\"https:\/\/arxiv.org\/pdf\/2508.00547v2\">https:\/\/arxiv.org\/pdf\/2508.00547v2<\/a><\/p>\n\n\n\n<p class=\"wp-block-paragraph\">[3] Dai et al. <em>Approximate Query Processing under Updates<\/em>. In SIGMOD 2026. <a href=\"https:\/\/doi.org\/10.1145\/3769760\">https:\/\/doi.org\/10.1145\/3769760<\/a><\/p>\n\n\n\n<p class=\"wp-block-paragraph\">[4] Niu et al. <em>MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing.<\/em> <a href=\"https:\/\/arxiv.org\/pdf\/2509.22186\">https:\/\/arxiv.org\/pdf\/2509.22186<\/a><\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><\/p>\n","protected":false},"excerpt":{"rendered":"<p>MinerU is an open-source document parser from OpenDataLab, the open-data group at Shanghai Artificial Intelligence Laboratory [1]. It converts PDFs, images and Microsoft Office files into Markdown, a plain-text format with simple marks for headings, lists and tables. It can also produce JSON that records the page layout. We installed MinerU 4.0 on a laptop [&hellip;]<\/p>\n","protected":false},"author":43,"featured_media":0,"template":"","meta":{"_acf_changed":false,"_ppma_block_editor_authors":""},"tags":[],"content_types":[{"id":50,"name":"Blog Post","slug":"article"}],"ppma_author":[{"id":43,"display_name":"Serafeim Papadias","first_name":"Serafeim","last_name":"Papadias","nickname":"serafeim.papadias","user_nicename":"serafeim-papadias","user_email":"serafeim.papadias@satalia.com","biographical_info":"Serafeim (Makis) Papadias is a Senior Data Scientist in Satalia's Research Lab, contributing to WPP Research's agenda on peer-to-peer communities of LLM agents. He holds a PhD from TU Berlin in scalable data systems and graph algorithms (VLDB, ICDE, EDBT) and previously built graph-based retrieval systems at Athena Research Center. He focuses on the network side of agent communities: how expertise is found, how teams form, and how information stays reliable as the population grows.","avatar_url":"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/makis_onboarding_cropped.jpeg","job_title":"Senior Data Scientist","is_lead":false,"display_as_researcher":true,"order_priority":null}],"class_list":["post-2137","research_feed","type-research_feed","status-publish","hentry","content_type-article"],"acf":{"content":"","content_quarter":"Q3 2026","related_pods":[1362]},"research_categories":[],"raw_acf":{"content":"","content_quarter":"Q3 2026","related_pods":["1362"],"featured":"","legacy_perspective_source_id":""},"_links":{"self":[{"href":"https:\/\/cms.research.wpp.com\/index.php?rest_route=\/wp\/v2\/research_feed\/2137","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/cms.research.wpp.com\/index.php?rest_route=\/wp\/v2\/research_feed"}],"about":[{"href":"https:\/\/cms.research.wpp.com\/index.php?rest_route=\/wp\/v2\/types\/research_feed"}],"author":[{"embeddable":true,"href":"https:\/\/cms.research.wpp.com\/index.php?rest_route=\/wp\/v2\/users\/43"}],"acf:post":[{"embeddable":true,"href":"https:\/\/cms.research.wpp.com\/index.php?rest_route=\/wp\/v2\/research_pods\/1362"}],"wp:attachment":[{"href":"https:\/\/cms.research.wpp.com\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=2137"}],"wp:term":[{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/cms.research.wpp.com\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=2137"},{"taxonomy":"content_type","embeddable":true,"href":"https:\/\/cms.research.wpp.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcontent_types&post=2137"},{"taxonomy":"author","embeddable":true,"href":"https:\/\/cms.research.wpp.com\/index.php?rest_route=%2Fwp%2Fv2%2Fppma_author&post=2137"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}