What is the fastest Python PDF library?

PDF Oxide is the fastest Python PDF library, with 0.8ms mean text extraction time — 5.8× faster than PyMuPDF (4.6ms) and 15× faster than pypdf (12.1ms). Benchmarked on 3,830 real-world PDFs with 100% pass rate.

Is PDF Oxide free for commercial use?

Yes. PDF Oxide is MIT licensed — free for all uses including commercial products, SaaS, and proprietary software. No license fees, no sales calls, no AGPL restrictions.

Can PDF Oxide handle scanned PDFs with OCR?

Yes. PDF Oxide includes built-in OCR via PaddleOCR and ONNX Runtime. No Tesseract installation needed — just pip install pdf_oxide and use extract_text_ocr(). Supports PP-OCRv3, v4, and v5 models.

Does PDF Oxide support XFA forms?

Yes. PDF Oxide is the only Python PDF library that can detect, analyze, and extract data from XFA forms (XML Forms Architecture). PyMuPDF, pypdf, pdfplumber, and pdfminer cannot read XFA form data.

How does PDF Oxide compare to PyMuPDF?

PDF Oxide is 5.8× faster than PyMuPDF (0.8ms vs 4.6ms mean), has a 100% pass rate vs 99.3%, and is MIT licensed vs PyMuPDF's AGPL-3.0. PDF Oxide also has built-in Markdown/HTML output and XFA form support that PyMuPDF lacks.

Can PDF Oxide convert PDF to Markdown?

Yes. PDF Oxide has built-in PDF to Markdown conversion with heading detection, table preservation, and list formatting — ideal for LLM and RAG pipelines. No separate package needed, unlike PyMuPDF which requires pymupdf4llm (69× slower).

Page API 레퍼런스

v0.3.34부터 모든 바인딩에서 Page 객체를 사용할 수 있습니다. 이제 모든 추출 호출에 page_index를 전달하는 대신 문서를 반복하며 각 페이지에서 직접 추출 메서드를 호출할 수 있습니다. 타입 이름은 Python, Node.js, C#, Go에서 일관되게 Page이며, Rust에서는 동일한 구조를 PdfPage로 제공합니다.

빠른 예제

Python

from pdf_oxide import PdfDocument

with PdfDocument("paper.pdf") as doc:
    for page in doc:                       # len(doc), doc[i], doc[-1] also work
        print(page.text[:80])
        md = page.markdown(detect_headings=True)

Rust

use pdf_oxide::api::Pdf;

let mut doc = Pdf::open("paper.pdf")?;
for i in 0..doc.page_count()? {
    let page = doc.page(i)?;
    println!("{}", &page.text()?[..80]);
}

JavaScript / TypeScript (Node)

const { PdfDocument } = require("pdf-oxide");

const doc = new PdfDocument("paper.pdf");
for (const page of doc) {
  console.log(page.extractText().slice(0, 80));
}
doc.close();

package main

import (
    "fmt"
    "log"
    pdfoxide "github.com/yfedoseev/pdf_oxide/go"
)

func main() {
    doc, err := pdfoxide.Open("paper.pdf")
    if err != nil { log.Fatal(err) }
    defer doc.Close()

    pages, _ := doc.Pages()
    for _, page := range pages {
        text, _ := page.ExtractText()
        fmt.Println(text[:80])
    }
}

using PdfOxide;

using var doc = PdfDocument.Open("paper.pdf");
foreach (var page in doc.Pages)
{
    Console.WriteLine(page.ExtractText()[..Math.Min(80, page.ExtractText().Length)]);
}

Java

import fyi.oxide.pdf.PdfDocument;
import java.nio.file.Path;

try (PdfDocument doc = PdfDocument.open(Path.of("paper.pdf"))) {
    for (int i = 0; i < doc.pageCount(); i++) {
        String text = doc.extractText(i);
        System.out.println(text.substring(0, Math.min(80, text.length())));
        String md = doc.toMarkdown(i);
    }
}

Kotlin

import fyi.oxide.pdf.PdfDocument

PdfDocument.open(java.nio.file.Path.of("paper.pdf")).use { doc ->
    for (i in 0 until doc.pageCount()) {
        val text = doc.extractText(i)
        println(text.substring(0, minOf(80, text.length)))
        val md = doc.toMarkdown(i)
    }
}

Scala

import fyi.oxide.pdf.PdfDocument
import scala.util.Using

Using.resource(PdfDocument.open("paper.pdf")) { doc =>
  for (i <- 0 until doc.pageCount()) {
    val text = doc.extractText(i)
    println(text.substring(0, math.min(80, text.length)))
    val md = doc.toMarkdown(i)
  }
}

Clojure

(require '[pdf-oxide.core :as pdf])

(with-open [doc (pdf/open "paper.pdf")]
  (doseq [i (range (pdf/page-count doc))]
    (let [text (pdf/extract-text doc i)]
      (println (subs text 0 (min 80 (count text))))
      (pdf/to-markdown doc i))))

Ruby

require 'pdf_oxide'

PdfOxide::PdfDocument.open('paper.pdf') do |doc|
  (0...doc.page_count).each do |i|
    text = doc.extract_text(i)
    puts text[0, 80]
    md = doc.to_markdown(i)
  end
end

PHP

use PdfOxide\PdfDocument;

$doc = PdfDocument::open('paper.pdf');
for ($i = 0; $i < $doc->pageCount(); $i++) {
    $text = $doc->extractText($i);
    echo substr($text, 0, 80), "\n";
    $md = $doc->toMarkdown($i);
}
$doc->close();

C++

#include <pdf_oxide/pdf_oxide.hpp>

auto doc = pdf_oxide::Document::open("paper.pdf");
for (int i = 0; i < doc.page_count(); i++) {
    auto text = doc.extract_text(i);
    std::cout << text.substr(0, 80) << "\n";
    auto md = doc.to_markdown(i);
}

Swift

import PdfOxide

let doc = try Document.open("paper.pdf")
for i in 0..<(try doc.pageCount()) {
    let text = try doc.extractText(i)
    print(text.prefix(80))
    let md = try doc.toMarkdown(i)
}

Dart

import 'package:pdf_oxide/pdf_oxide.dart';

final doc = PdfDocument.open('paper.pdf');
for (var i = 0; i < doc.pageCount; i++) {
  final text = doc.extractText(i);
  print(text.substring(0, text.length < 80 ? text.length : 80));
  final md = doc.toMarkdown(i);
}
doc.close();

library(pdfoxide)

doc <- pdf_open("paper.pdf")
for (i in 0:(pdf_page_count(doc) - 1)) {
  text <- pdf_extract_text(doc, i)
  cat(substr(text, 1, 80), "\n")
  md <- pdf_to_markdown(doc, i)
}

Julia

using PdfOxide

doc = open_document("paper.pdf")
for i in 0:(page_count(doc) - 1)
    text = extract_text(doc, i)
    println(first(text, 80))
    md = to_markdown(doc, i)
end

Zig

const pdf_oxide = @import("pdf_oxide");
const a = std.heap.page_allocator;

var doc = try pdf_oxide.Document.open("paper.pdf");
var i: usize = 0;
while (i < try doc.pageCount()) : (i += 1) {
    const text = try doc.extractText(a, i);
    std.debug.print("{s}\n", .{text[0..@min(80, text.len)]});
    const md = try doc.toMarkdown(a, i);
}

Objective-C

#import "POXPdfOxide.h"
NSError *err = nil;

POXDocument *doc = [POXDocument openPath:@"paper.pdf" error:&err];
for (NSInteger i = 0; i < [doc pageCountError:&err]; i++) {
    NSString *text = [doc extractText:i error:&err];
    NSLog(@"%@", [text substringToIndex:MIN(80, text.length)]);
    NSString *md = [doc toMarkdown:i error:&err];
}

Elixir

{:ok, doc} = PdfOxide.open("paper.pdf")
{:ok, n} = PdfOxide.page_count(doc)
for i <- 0..(n - 1) do
  {:ok, text} = PdfOxide.extract_text(doc, i)
  IO.puts(String.slice(text, 0, 80))
  {:ok, md} = PdfOxide.to_markdown(doc, i)
end

Python — `Page`

지연 속성 방식 — 콘텐츠는 첫 번째 접근 시 파싱되어 Page에 캐시됩니다.

멤버	반환값	설명
`page.text`	`str`	추출된 텍스트 (열 인식)
`page.chars`	`list[Char]`	bbox 및 폰트 정보가 포함된 문자 단위 레코드
`page.words`	`list[Word]`	bbox가 포함된 단어 단위 레코드
`page.lines`	`list[TextLine]`	bbox가 포함된 텍스트 줄
`page.spans`	`list[Span]`	스타일이 적용된 스팬 (폰트, 크기, 굵기)
`page.tables`	`list[Table]`	구조화된 테이블 행과 셀 bbox
`page.images`	`list[Image]`	이미지 메타데이터
`page.paths`	`list[Path]`	벡터 패스 레코드
`page.annotations`	`list[Annotation]`	해당 페이지의 주석
`page.markdown(detect_headings=True)`	`str`	Markdown 변환
`page.plain_text()`	`str`	일반 텍스트 (레이아웃 힌트 없음)
`page.html()`	`str`	HTML 변환
`page.render(format="png")`	`bytes`	페이지를 PNG / JPEG로 렌더링
`page.search(term, case_sensitive=False)`	`list[SearchResult]`	해당 페이지에서 텍스트 검색
`page.region(rect)`	`PageRegion`	지정 영역 내 범위 추출

with PdfDocument("paper.pdf") as doc:
    page = doc[0]                 # or doc.page(0)
    for word in page.words:       # first access parses; subsequent calls cached
        print(word.text, word.bbox)

    # Scoped extraction
    header = page.region((0, 700, 612, 92)).extract_text()

기존의 쓰기용 편집기 클래스 PdfPage는 변경되지 않았습니다. 새로운 Page는 완전히 읽기 전용입니다.

Rust — `PdfPage`

use pdf_oxide::api::Pdf;

let mut doc = Pdf::open("paper.pdf")?;
let page = doc.page(0)?;

let text = page.text()?;
let words = page.extract_words()?;
let tables = page.extract_tables()?;
let md = page.to_markdown(true)?;

PdfPage에서 사용 가능한 메서드:

text(), plain_text(), to_markdown(detect_headings), to_html()
extract_chars(), extract_words(), extract_lines(), extract_spans()
extract_tables(), extract_paths(), extract_images()
annotations(), render(format)
search(term) — 범위 내 검색
find_text_containing(substring) — ID 포함 DOM 수준 일치 목록

Node.js — `Page`

const { PdfDocument } = require("pdf-oxide");

const doc = new PdfDocument("paper.pdf");
const page = doc.page(0);

console.log(page.width, page.height, page.rotation);  // cached
console.log(page.extractText());
const words = page.extractWords();
const tables = page.extractTables();
const md = page.toMarkdown();

PdfDocument는 Symbol.iterator를 통한 for..of를 지원하며, doc.page(i)와 doc.pageCount()도 사용할 수 있습니다.

기존에 네이티브 전용이었던 6개 메서드가 TS 레이어를 통해 Page와 PdfDocument 모두에서 사용 가능해졌습니다:

extractWords
extractTextLines
extractTables
extractPaths
getEmbeddedImages
ocrExtractText

각 메서드에는 비동기 버전(extractTextAsync, toMarkdownAsync 등)이 있습니다.

Go — `Page`

doc, _ := pdfoxide.Open("paper.pdf")
defer doc.Close()

page, _ := doc.Page(0)
text, _ := page.ExtractText()
md, _   := page.ToMarkdown()
tables, _ := page.ExtractTables()

// Iterate every page
all, _ := doc.Pages()
for i, p := range all {
    t, _ := p.ExtractText()
    fmt.Printf("page %d: %d chars\n", i, len(t))
}

Go의 Page 구조체는 전체 메서드를 제공합니다: ExtractText, ToMarkdown, ToHtml, ToPlainText, ExtractWords, ExtractTextLines, ExtractTables, ExtractChars, ExtractPaths, Annotations, Images, Fonts, RenderPage, Search.

C# — `Page`

using PdfOxide;

using var doc = PdfDocument.Open("paper.pdf");

Page page = doc[0];                            // or doc.Pages[0] or doc.Page(0)
string text = page.ExtractText();
string md   = page.ToMarkdown();
Table[] tables = page.ExtractTables();

// Async variants
string textAsync = await page.ExtractTextAsync();
string mdAsync   = await page.ToMarkdownAsync();

doc.Pages는 IReadOnlyList<Page>입니다. 모든 동기 메서드에는 CancellationToken을 지원하는 async Task<T> 대응 버전이 있습니다.

구조화된 테이블 형식

extract_tables()(PdfDocument와 Page 모두에서 사용 가능)는 언어 전반에 걸쳐 일관된 Table 타입을 반환합니다:

언어	타입	셀 접근 방식
Rust	`Table`	`rows[i].cells[j]` 반복
Python	`dict`	`row["cells"][i]["text"]`
Go	`Table`	`table.CellText(row, col)`
C#	`Table`	`table.CellText(row, col)`
Node.js	`Table` 인터페이스	`table.cells[row][col]`

각 셀에는 텍스트와 경계 박스가 포함되어 있어 추출 결과를 페이지 좌표와 연결할 수 있습니다.

`doc.extract_*(page_index)` 에서 마이그레이션

이전 방식 (계속 지원):

doc = PdfDocument("paper.pdf")
for i in range(doc.page_count()):
    print(doc.extract_text(i))
    print(doc.to_markdown(i, detect_headings=True))
    print(doc.extract_tables(i))

새 방식 (v0.3.34+):

with PdfDocument("paper.pdf") as doc:
    for page in doc:
        print(page.text)
        print(page.markdown(detect_headings=True))
        print(page.tables)

두 방식 모두 계속 지원됩니다. Page 방식은 페이지 단위 파이프라인에서 가독성이 높고, 반복적인 인덱스 관리를 없애줍니다.

Page API 레퍼런스

빠른 예제

Python — Page

Rust — PdfPage

Node.js — Page

Go — Page

C# — Page