Chonkie
A Python ingestion library that chunks documents, refines passages, adds embeddings and connects the results to vector databases for RAG pipelines.
About Chonkie
Developers need reusable document chunking and ingestion components rather than rebuilding them for every retrieval application.
Who it’s for
- Developers building document retrieval pipelines
- Teams preparing text for embeddings and vector databases
When to consider it
Consider Chonkie when you want to compose document chunking, refinement and embedding steps in a Python pipeline.
Tradeoffs & limitations
- Semantic, neural and provider-specific features require optional dependencies beyond the base installation.
- The library is MIT licensed; deploying its self-hosted API requires running and maintaining the server.
SOTA overview · Documentation-based assessment · Sources & review method
Updates
No updates shared yet.
Discussion
Newest firstAsk a question or share how you use Chonkie.
Keep it helpful. Community rules
Loading discussion…
Sources & review method
Documentation-based assessment · Oct 8, 2026 · Prepared with AI assistance; not a hands-on benchmark.
- Chonkie canonical repository and pipeline examples
- Chonkie MIT license
- Chonkie optional installation dependencies
- Chonkie recent commits
- Feyn organization and owner domain
Checked by SOTA · AI-assisted documentation review. Selection advice is our assessment; verify current requirements for your deployment.
Import history & original evidence
Chonkie on GitHub
A Python ingestion library that chunks documents, refines passages, adds embeddings and connects the results to vector databases for RAG pipelines.
Based on official pages and announcements checked on 2026-10-08. No hands-on test, performance benchmark or popularity ranking is claimed.