Metadata-Version: 2.4
Name: xml2table
Version: 0.1.0
Summary: Convert XML and PDF documents to CSV, Excel, XML, and text with a small, predictable SDK.
Author-email: Meet2147 <meetjethwa3@gmail.com>
License: MIT
Project-URL: Homepage, https://github.com/Meet2147/pythonLibraries/tree/main/xml2table
Project-URL: Repository, https://github.com/Meet2147/pythonLibraries
Project-URL: Issues, https://github.com/Meet2147/pythonLibraries/issues
Keywords: xml,pdf,csv,excel,xlsx,convert,flatten,etl
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Developers
Classifier: License :: OSI Approved :: MIT License
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.8
Classifier: Programming Language :: Python :: 3.9
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Topic :: Software Development :: Libraries
Classifier: Topic :: Text Processing :: Markup :: XML
Requires-Python: >=3.8
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: openpyxl>=3.1
Provides-Extra: pandas
Requires-Dist: pandas>=1.3; extra == "pandas"
Provides-Extra: lxml
Requires-Dist: lxml>=4.9; extra == "lxml"
Provides-Extra: pdf-text
Requires-Dist: pypdf>=4.0; extra == "pdf-text"
Provides-Extra: pdf
Requires-Dist: pdfplumber>=0.10; extra == "pdf"
Requires-Dist: pypdf>=4.0; extra == "pdf"
Provides-Extra: dev
Requires-Dist: pytest>=7.0; extra == "dev"
Requires-Dist: pandas>=1.3; extra == "dev"
Requires-Dist: pdfplumber>=0.10; extra == "dev"
Requires-Dist: pypdf>=4.0; extra == "dev"
Requires-Dist: fpdf2>=2.7; extra == "dev"
Dynamic: license-file

# xml2table

A small, predictable Python SDK for converting XML and PDF documents to
CSV, Excel, XML, and plain text.

- **XML → CSV / Excel.** One core recursive flattening function turns
  nested XML into flat rows; CSV and Excel are thin writers on top of it.
  No data loss by default: repeated elements can be joined into one cell,
  exploded into extra rows, or spread across indexed columns — you choose.
- **PDF → XML.** A structure-preserving extractor ([`pdfplumber`](https://github.com/jsvine/pdfplumber)-backed)
  groups a PDF's words into paragraphs and detects tables, keeping both in
  their original top-to-bottom reading order — no data lost, no tables
  flattened into loose words.
- **PDF → text.** A fast, lightweight raw-text extractor
  ([`pypdf`](https://github.com/py-pdf/pypdf)-backed) that doesn't need
  table detection at all — just the PDF's text, in reading order, with its
  visual whitespace layout preserved by default.
- **Small dependency footprint.** Only [`openpyxl`](https://openpyxl.readthedocs.io/)
  is required for the XML-to-table side. `pandas` (DataFrames), `pypdf`
  (PDF text), and `pdfplumber` (PDF XML/tables) are all optional extras —
  importing `xml2table` never requires any of them.
- **Three ways in, for each conversion:** one-line functions, a reusable
  converter object, or the `xml2table` CLI.

## Install

```bash
pip install -e .             # from a checkout: XML -> CSV/Excel only
pip install -e ".[pandas]"   # + optional DataFrame support
pip install -e ".[pdf-text]" # + PDF -> text (pypdf only, lightweight)
pip install -e ".[pdf]"      # + PDF -> XML and text (pdfplumber + pypdf)
```

## Quick start

```python
from xml2table import xml_to_csv, xml_to_excel

xml_to_csv("orders.xml", "orders.csv")
xml_to_excel("orders.xml", "orders.xlsx")
```

Given:

```xml
<orders>
  <order id="1">
    <customer><name>Jane Doe</name></customer>
    <total>99.99</total>
  </order>
</orders>
```

you get one row per `<order>`:

| @id | customer.name | total |
|-----|----------------|-------|
| 1   | Jane Doe       | 99.99 |

Nested elements are flattened with `.`-separated keys, attributes get an
`@` prefix, and the record element (`order` here) is auto-detected as "the
repeated child of the root". When detection is ambiguous, pass
`record_path` explicitly.

## Reusable converter

Parse once, write many times:

```python
from xml2table import XMLConverter, FlattenOptions

converter = XMLConverter("orders.xml", options=FlattenOptions(record_path="Order"))
converter.to_csv("orders.csv")
converter.to_excel("orders.xlsx", sheet_name="Orders")
rows = converter.to_records()      # list[dict]
df = converter.to_dataframe()      # requires pandas
```

## Handling repeated elements (`array_mode`)

Given an order with two line items, `FlattenOptions.array_mode` controls the
shape of the output:

| mode | Result | Use when |
|------|--------|----------|
| `"join"` (default) | One row per order; items collapsed into one joined cell | You just want a quick, human-readable table |
| `"explode"` | One row per item (order fields repeat) | You want a normalized, analysis-ready table, like a SQL join |
| `"index"` | One row per order; items spread into `items.item.0.*`, `items.item.1.*`, ... | You need every field as its own column with no row duplication |

```python
from xml2table import FlattenOptions, xml_to_records

xml_to_records("orders.xml", options=FlattenOptions(array_mode="explode"))
```

## Multi-sheet Excel from one document

Turn different parts of the same XML document into separate, related sheets
(e.g. an "orders" table and an "items" table):

```python
from xml2table import xml_to_excel

xml_to_excel(
    "orders.xml", "orders.xlsx",
    sheets={"orders": "Order", "items": ".//Order/Items/Item"},
)
```

## Record paths

`record_path` uses the same syntax as
[`Element.findall`](https://docs.python.org/3/library/xml.etree.elementtree.html#xml.etree.ElementTree.Element.findall)
(a subset of XPath): `"Order"`, `"Orders/Order"`, `".//Item"`,
`"Order[@status='shipped']"`, etc.

## CLI (XML)

```bash
xml2table csv orders.xml orders.csv --record-path Order --array-mode explode
xml2table excel orders.xml orders.xlsx --sheet-name Orders
xml2table excel orders.xml orders.xlsx --sheet orders=Order --sheet items=".//Item"
```

Run `xml2table csv --help` or `xml2table excel --help` for all flags
(`--separator`, `--attribute-prefix`, `--no-attributes`, `--keep-namespaces`,
`--delimiter`, `--encoding`, ...).

## PDF → XML / text

`pdf_to_xml` and `pdf_to_text` are two independent, purpose-built backends
behind one options object and one CLI command:

- **`pdf_to_xml`** (pdfplumber) groups the PDF's words into paragraphs (by
  line, then by vertical gap) and detects tables separately via ruling
  lines or text alignment, then places both back in the order they appear
  on the page. Nothing is dropped, and table text never bleeds into
  surrounding paragraphs.
- **`pdf_to_text`** (pypdf) is a much lighter path: it just extracts each
  page's text in reading order, preserving the PDF's visual whitespace
  layout by default (so simple tables and columns still read naturally)
  without doing any table detection.

```python
from xml2table import pdf_to_xml, pdf_to_text

pdf_to_xml("invoice.pdf", "invoice.xml")    # paragraphs + tables (needs xml2table[pdf])
pdf_to_text("invoice.pdf", "invoice.txt")   # fast raw text (needs xml2table[pdf-text])
```

`invoice.xml` looks like:

```xml
<document source="invoice.pdf" pages="1">
  <page number="1" width="595.28" height="841.89">
    <paragraph bbox="42.83,43.13,159.90,61.13">Invoice #1024</paragraph>
    <table bbox="40.00,220.00,540.00,316.00" rows="4" cols="3">
      <row><cell>Item</cell><cell>Qty</cell><cell>Price</cell></row>
      <row><cell>Widget</cell><cell>2</cell><cell>$10.00</cell></row>
      ...
    </table>
  </page>
</document>
```

`invoice.txt` is pypdf's layout-preserving text, e.g.:

```
Invoice #1024

Bill To: Jane Doe
123 Example Street
Springfield, USA

Thank you for your business. Payment is due within thirty days...

Item                                                          Qty        Price
Widget                                                        2          $10.00
...
```

Reuse one `PDFConverter` for both (it lazily parses with pdfplumber only if
you call `to_xml()`/`pages`, and always uses pypdf for `to_text()`):

```python
from xml2table import PDFConverter

converter = PDFConverter("invoice.pdf")
converter.to_xml("invoice.xml")
converter.to_text("invoice.txt")
for page in converter.pages:
    print(page.number, len(page.paragraphs), len(page.tables))
```

### `PDFOptions` reference

| Option | Default | Used by | Description |
|--------|---------|---------|--------------|
| `line_tolerance` | `3.0` | XML | Max vertical gap (points) for words to count as the same line |
| `paragraph_gap` | `6.0` | XML | Min vertical gap (points) between lines that starts a new paragraph |
| `table_settings` | `None` | XML | Passed through to pdfplumber's `find_tables()` for unusual tables (e.g. borderless). **Caution:** this applies to the whole page, not just the table — see [Validation](#validation-tested-on-1200-financial-report-pdfs) below before using it on documents with narrative text. |
| `cell_na` | `""` | XML | String used for empty/missing table cells |
| `page_separator` | `"\n----- Page {page} -----\n"` | text | Inserted between pages |
| `keep_layout` | `True` | text | Preserve the PDF's whitespace layout (pypdf `"layout"` mode) vs. plain, whitespace-normalized text |

### CLI (PDF)

```bash
xml2table pdf invoice.pdf invoice.xml --to xml
xml2table pdf invoice.pdf invoice.txt --to text
```

Run `xml2table pdf --help` for all flags (`--line-tolerance`,
`--paragraph-gap`, `--cell-na`, `--page-separator`, `--no-layout`).

## Validation: tested on 1,200 financial-report PDFs

`pdf_to_xml` and `pdf_to_text` were benchmarked against 1,200 generated
financial-report PDFs (balance sheets, income statements, cash-flow
statements, MD&A-style narrative text, footnotes — real financial
formatting: `$1,234,567`, `(123,456)` negatives, `N/A` blanks, multi-page,
ruled and borderless tables) with **exact ground truth** for every paragraph
and table cell, so the numbers below are measured, not estimated. Full
methodology, the generator, and raw per-document results are in
[`benchmarks/RESULTS.md`](benchmarks/RESULTS.md).

| Metric | Result |
|---|---|
| Documents converted without error | **1,200 / 1,200 (100%)** |
| Paragraph text fidelity (XML) | **100.000%** |
| Ruled-table shape + cell fidelity | **100.000%** |
| Borderless-table shape detection | 0.000% (documented limitation — see below) |
| Raw-text content recall (`pdf_to_text`) | **100.000%**, including for borderless tables |
| Throughput | 20.4 PDFs/sec (`pdf_to_xml`), 184.6 PDFs/sec (`pdf_to_text`) |

**The one real limitation, quantified:** pdfplumber's default table finder
needs ruling lines, so it doesn't detect borderless (text-only-aligned)
tables — but nothing is lost when it doesn't: the un-detected table's text
still comes through as ordinary paragraph text (`pdf_to_text` recall stays
at 100%). We also tested the obvious "fix" — pdfplumber's `table_settings`
override for borderless tables — across all 1,200 documents, and it made
things *worse*: because the override applies to the whole page, it started
misreading ordinary paragraph sentences as table cells, corrupting
paragraph output in **97.9% of documents** (paragraph fidelity dropped from
100% to 13.9%). We did not ship that as a recommended workaround; see
`PDFOptions.table_settings`'s docstring and `benchmarks/RESULTS.md` for the
full numbers and why.

We could not download real financial filings for this test — this sandboxed
session's network policy blocks direct access to sites like sec.gov — so
the corpus is synthetic but built to real financial-statement conventions
specifically so every value has a known-correct answer to grade against.
`benchmarks/RESULTS.md` explains this in more detail and gives the exact
commands to reproduce or extend the benchmark (e.g. against real filings, on
a machine with broader network access).

## `FlattenOptions` reference

| Option | Default | Description |
|--------|---------|--------------|
| `record_path` | `None` (auto-detect) | Path to the repeated record element |
| `attribute_prefix` | `"@"` | Prefix for attribute-derived columns |
| `text_key` | `"#text"` | Key for an element's own text when it also has attributes/children |
| `separator` | `"."` | Separator for nested key paths |
| `array_mode` | `"join"` | `"join"`, `"explode"`, or `"index"` |
| `join_separator` | `"; "` | Separator used by `"join"` mode |
| `include_attributes` | `True` | Include XML attributes as columns |
| `strip_namespaces` | `True` | Strip `{namespace}` from tag/attribute names |
| `encoding` | `"utf-8"` | Text encoding for reads/writes |

## Errors

All exceptions inherit from `xml2table.XMLConversionError`:

- `XMLParseError` — malformed XML input
- `RecordPathNotFoundError` — `record_path` matched nothing, or automatic
  record detection was ambiguous (the error message tells you what to pass)
- `PDFExtractionError` — a PDF file could not be opened or parsed
- `MissingOptionalDependencyError` — e.g. calling `to_dataframe()` without
  `pandas`, `pdf_to_xml()` without `pdfplumber`, or `pdf_to_text()` without
  `pypdf`, installed

## Development

```bash
pip install -e ".[dev]"   # includes pandas, pdfplumber, pypdf, and fpdf2 (for regenerating PDF fixtures)
pytest
python examples/quickstart.py
```

Project layout:

```
src/xml2table/
  parser.py          # XML -> list[dict] flattening engine
  options.py         # FlattenOptions
  converter.py       # XMLConverter + module-level convenience functions
  writers/           # CSV and Excel output
  pdf_extract.py     # pdfplumber: PDF -> Document(pages of Paragraph/Table), in reading order
  pdf_xml_writer.py  # Document -> XML
  pdf_text_extract.py # pypdf: PDF -> plain text, independent of pdf_extract.py
  pdf_options.py     # PDFOptions
  pdf_converter.py   # PDFConverter + pdf_to_xml/pdf_to_text
  cli.py             # `xml2table` command-line tool
tests/
  fixtures/          # sample XML documents and PDF fixtures (fixtures/pdf/)
  test_*.py
examples/
  quickstart.py
```
