Unified API
One OpenAI-compatible endpoint for /v1/chat/completions, /v1/completions and /v1/embeddings, streamed over SSE. Keep your SDK; change the base URL.
Token Gear by FreshToken · LLM token aggregation & redistribution
Token Gear aggregates model capacity from many NCP Token Engines and redistributes it to tenants through one OpenAI-compatible API — with session-aware routing, SLA and security policy, quotas, and prepaid metering to the token.
https://api.freshtoken.ai/v1
Why Token Gear
Token Gear owns business-level identity, routing and policy, so tenants get one reliable API and engines get steady, well-shaped demand.
One OpenAI-compatible endpoint for /v1/chat/completions, /v1/completions and /v1/embeddings, streamed over SSE. Keep your SDK; change the base URL.
cheapest, fastest, balanced or pinned, on live health with circuit breakers and failover. Each session stays sticky to one Token Engine.
Organizations as tenants, with per-key quotas, spend caps, rate limits and model allow-lists. Keys are shown once and stored as hashes.
The most a request can cost is reserved before it goes upstream and settled to the token when it ends. Shared org balances, monthly invoices, reseller referrals.
Architecture
Token Gear sits between tenants and NCP domains. It keeps the logical session; each Token Engine keeps the execution session. A bidirectional API joins the two.
Individual users
Tenants · business users
Aggregator / Redistribution Platform · session identity, routing, SLA and security policy
NCP A
NCP B
NCP C
Five principles
Sharing is what makes redistribution efficient. These five rules decide where it goes, and where it stops.
Sharing is optimized across security, efficiency and flexibility — not maximized indiscriminately.
One Token Engine can serve sessions from multiple tenants at the same time.
Sessions are allocated over time; capacity is reallocated to the next session when one completes.
Token Gear manages logical sessions at the tenant level; the Token Engine manages execution sessions at the engine level.
Each session is sticky to one Token Engine for its lifetime. No cross-engine migration within a session.
Featured models
Open-weight models served by NCP Token Engines, alongside frontier models from provider APIs.
DeepSeek
Moonshot AI
Qwen
Meta
OpenAI
Anthropic
Two sides, one platform
Tenants and developers buy tokens through one API. NCPs running Token Engines supply capacity and receive redistributed demand.
For builders & tenants
For NCPs / Token Engine operators
Get started
Every public endpoint lives under /v1.
Create an account, personal or for your organization.
Top up a prepaid balance; each request is metered to the token.
Create a key starting with ft-. It is shown once.
curl https://api.freshtoken.ai/v1/chat/completions \
-H "Authorization: Bearer ft-..."
\
-H "Content-Type: application/json"
\
-d '{
"model": "deepseek/deepseek-v3",
"messages": [{"role": "user", "content": "Hello, Token Gear"}],
"stream": true
}'
from openai import OpenAI
client = OpenAI(
base_url="https://api.freshtoken.ai/v1"
,
api_key="ft-..."
,
)
reply = client.chat.completions.create(
model="deepseek/deepseek-v3"
,
messages=[{"role"
: "user"
, "content"
: "Hello, Token Gear"
}],
)
print(reply.choices[0].message.content)import OpenAI from "openai"
;
const client = new OpenAI({
baseURL: "https://api.freshtoken.ai/v1"
,
apiKey: process.env.TOKEN_GEAR_API_KEY, // ft-...
});
const reply = await client.chat.completions.create({
model: "deepseek/deepseek-v3"
,
messages: [{ role: "user"
, content: "Hello, Token Gear"
}],
});
console.log(reply.choices[0].message.content);FreshToken's LLM token aggregation and redistribution platform. It aggregates capacity from many Token Engines and provider APIs, and redistributes it to tenants through one OpenAI-compatible API, handling session identity, routing, SLA and security policy along the way.
You hold one key and one balance instead of an account per provider. Token Gear routes each request on live health with failover, applies your tenant's quotas and policies, and meters every request to the token against a single prepaid balance.
An NCP is a compute provider domain. Each NCP runs a Token Engine: execution session state and an inference scheduler (P/D disaggregation, KV cache, batching, DP/EP) on top of its LLM hosting and GPU infrastructure.
Sessions have two layers. Token Gear keeps the logical session — tenant identity, tracking and quotas. The Token Engine keeps the execution session. A session stays on one Token Engine for its lifetime and is never migrated between engines midway.
Prepaid credits, metered per token. The most a request can cost is reserved before it goes upstream, and the actual tokens are settled when it ends. Organizations share a balance, and statements are issued monthly.
Your content is passed to the engine or provider that serves the request and back to you. We keep what is needed to run and bill an account — keys, usage records and balances — not your prompts or responses. See the privacy policy.
Yes. Chat completions, completions and embeddings follow the OpenAI format, including SSE streaming. Point the SDK's base URL at Token Gear and use an ft- key.
Write to support@freshtoken.ai. We connect your Token Engine through the bidirectional API, register the models it hosts, and start redistributing tenant demand to it.
One API, one balance, sessions routed on live health.