Rare find

Textract. Pull text out of any document automatically.

textract.readthedocs.io

Vaya's read on this project

Problem, audience, market, and the verdict — sign in to see it.

Updates

July 2026
  • ci: resolve CI failures (#587)
  • test(#374): verify that cx-freeze and similar do not break PATH usage…
  • feat: Drop Python 3.9 (#585)
  • fix(#335): more precise error messages (#575)
  • feat(#308): add .ods Spreadsheet support (#583)
  • feat(#51): replace deprecated antiword with LibreOffice (#582)
  • test: resolve false positive warnings & add optional pre-commit (#584)
  • fix: prevent opening a cmd window with Popen() on Windows (#574)
  • fix(#544): modernize xlsx handling (#572)
May 2026
  • fix(#185): resolve missing wheel (#560)
  • test(#168): expand tests of unbalanced parenthesis and similar (#562)
  • fix(#224): remove unused statement (#559)
  • ci: One PR per month with auto-updates (#561)
  • Bump lxml from 6.0.2 to 6.1.0 (#556)
  • Bump cryptography from 46.0.6 to 46.0.7 (#557)
  • Bump softprops/action-gh-release from 2 to 3 in the all-dependencies …
April 2026
  • Release —v2.0.0
  • Release —v2.0.0rc1
  • feat: migrate to uv, expand tests, CI/CD, and modernizations from tex…
February 2026
  • fix(#342): return null on empty stream (#422)
  • fix: Enable encoding detection for the txt parser (#456)
  • fix: catch ShellError from pdf2txt.py (#495)
August 2021
  • Release —v1.6.4
July 2019
  • Release —v1.6.3
  • Release —v1.6.2
June 2017
  • Release —v1.6.1
April 2017
  • Release —v1.6.0
November 2016
  • Release —v1.5.0
October 2015
  • Release —v1.4.0
June 2015
  • Release —v1.3.0