Claupt
✂️
AI Data Preprocessing Tool

Smart Text Chunker & Slicer

Split massive documents into token-sized chunks while preserving paragraph, sentence, or word boundaries. Designed for AI context windows, RAG pipelines, prompt preparation, and large-document preprocessing.

📝 Your Text

Paste text or load a local document.

Chars: 0 Words: 0 Est. tokens: 0 Paragraphs: 0

⚙️ Chunking Settings

Presets:
✂️

Paste or upload a document to see your chunks here.

Split text without losing context

Choose a target size, unit and boundary mode. Preview chunks before downloading text or ZIP. Paragraph and sentence modes keep those units intact when possible; a long unit can exceed the target.

Example: for a long article, start with paragraph boundaries and inspect oversized chunks. Word targets use actual whitespace-separated word counts. Hard mode may split words but preserves Unicode characters.

Token targets are approximate. Overlap repeats part of the previous chunk and increases size; headers also add text. Allow room for system instructions and the response. Verify each final chunk against your provider’s tokenizer.