Home Projects docstrange
docstrange

docstrange

by NanoNets · GitHub

Extract and convert data from any document, images, pdfs, word doc, ppt or URL into multiple formats (Markdown, JSON, CSV, HTML) with intelligent structured data extraction and advanced OCR.

View on GitHub
⭐ Stars
1.5k
🍴 Forks
134
📜 License
MIT
Commercial use OK
📅 Created
2025
🔄 Last commit
8 mo ago
🏷️ Category
ai
💻 Language
You maintain this project?

Claim its page: indexed whatever its rank, translated into six languages, and enriched with what you write yourself.

Claim this page →
docstrange — GitHub preview card
📈 Star history
1 5041 503
2026-07-202026-07-21
📄 About

Extract and convert data from any document, images, pdfs, word doc, ppt or URL into multiple formats (Markdown, JSON, CSV, HTML) with intelligent structured data extraction and advanced OCR.

Frequently asked questions

What is docstrange?

Extract and convert data from any document, images, pdfs, word doc, ppt or URL into multiple formats (Markdown, JSON, CSV, HTML) with intelligent structured data extraction and advanced OCR.

Is docstrange open source?

docstrange is an open-source project. It is released under the MIT license.

Is docstrange free?

Yes. docstrange is free and open source — you can use, modify and self-host it.

🏅 Maintainer of this project?
OpenSourceAI badge — docstrange

Add this live badge to your README — your GitHub stars and directory rank, refreshed daily.

[![OpenSourceAI](https://opensourceai.tech/badge.php?tool=nanonets-docstrange)](https://opensourceai.tech/project/nanonets-docstrange.html)
More badge options →
🧬 Related projects🧬 View the DNA map →