Token Cut: Intelligent LLM Model Routing
Token Cut is a service that reduces LLM token costs by routing each request through Jev, a TypeSafe decision model that classifies the task and selects the lowest-priced eligible model within your configured constraints. It provides a single API for text, image, video, and audio generation across multiple providers.
Key Features
- Jev Model Routing: A small decision model (System One) that answers bounded questions (Choice, Score, Noul) to classify task capability needs.
- Cost Optimization: Automatically routes to the cheapest model that meets capability, modality, context, provider, and budget requirements.
- Unified API: One API key for four providers (OpenRouter, kie.ai, Replicate, fal.ai) covering text, image, video, and audio.
- Budget Controls: Set budget limits, capability matching, provider selection, and transparent billing.
- Savings Calculator: Estimate potential savings based on monthly LLM spend and expected provider savings.
- Use Cases: Email/support triage, document classification, model routing, and gating review steps.
How It Works
- Understand: Jev classifies the task and estimates required capability.
- Match & Route: Filter by modality, context, and budget; choose the lowest-priced eligible route.
- Measure & Account: Record model, provider, routing confidence, usage, and cost.
Token Cut is ideal for developers and businesses looking to cut LLM costs without sacrificing quality, offering a transparent and automated routing layer.

