pdfplumber

pdfplumber

pdfplumber is a Python library that provides detailed information about each text character, rectangle, and line in a PDF.

FreeCLIPython LibraryOpen sourceKorean
Visit websitegithub.com
Compare with book-to-skillExplore pdfplumber alternatives

Overview

pdfplumber is a Python library that provides detailed information about each text character, rectangle, and line in a PDF. Built on top of pdfminer.six, it excels at table extraction and visual debugging, making it a favorite for data journalists and analysts who need precise control over PDF data extraction and layout analysis.

Key features

  • Precise text coordinate extraction
  • Robust table extraction
  • Visual debugging support
  • Image and shape analysis
  • PDF metadata access
  • Custom extraction logic
  • Page layout visualization

Pricing

FreeStarting price: Free (open source)
View pricing page

Verified on:

Use cases

  • Automating table extraction from reports
  • Document layout analysis
  • Data extraction for journalism
  • PDF text cleaning

Who it is for

Python DevelopersData AnalystsData Journalists

Integrations

PandasJupyter Notebookpdfminer.sixPillow

How we verified this

Company, pricing, and feature details come from the primary sources below and our latest verification pass. When sources disagree, the official source and the most recent check win.

Last verified 08/30/2026Verified sources: 1

Alternatives

Tools you can use instead