The heading levels a PDF does not have
Every heading in the converted PDF came out at the same level. That is not a bug in your file, it is what the layout model can and cannot see.
for item in document.texts:
if item.label == "section_header":
print(item.level, item.text)Four headings, all level 1, although Returns is set in twenty point bold and the other three in fourteen. The layout model's job is to say this is a heading. It is not asked how deep the heading sits, so everything comes back flat.
Turning the levels on
Docling can work the depth out afterwards from three signals: the PDF's own bookmarks, numbering like 2.1 at the start of a heading, and the font each heading is set in. It is off by default because it is guesswork, and you ask for it with an options object.
from docling.datamodel.pipeline_options import HeadingHierarchyOptions
options = PdfPipelineOptions()
options.heading_hierarchy_options = HeadingHierarchyOptions(enabled=True)
options.generate_parsed_pages = True
deep = DocumentConverter(
format_options={InputFormat.PDF: PdfFormatOption(pipeline_options=options)})
nested = deep.convert("returns.pdf").document
for item in nested.texts:
if item.label == "section_header":
print(item.level, item.text)Now Returns is level 1 and the three smaller headings are level 2, recovered from nothing but the point size. The second line matters as much as the first: generate_parsed_pages is what keeps the font information around, and without it the font signal has nothing to read.
What it changes downstream
print(nested.export_to_markdown())Three hashes where there were two. The bigger effect arrives in part 6: a chunk records the headings above it, and with a flat document every heading looks like a top level one, so a chunk from deep inside a report is labelled as if it were a chapter.
Formats that never needed this
Markdown, HTML and Word files say the depth in the file, so their levels were right from lesson 6 without anyone asking. This option exists only because a PDF throws that information away.
- Set every heading in
make_returns_pdf.pyto the same point size and convert again. - Turn the option on but leave
generate_parsed_pagesoff, and see what you get. - Export the nested document as DocTags and find the level in the tag names.
Little by little, you're building something great.