Inference provider · Haifa, Israel
Inference costs less on hardware built for it.
Tannit AI is a new company serving leading open models through an OpenAI-compatible API. The hardware and the inference stack are our own — engineered around energy efficiency, which is why our cost per token comes in below the market rate for the same model.
What we do
Three things, plainly
- We build the hardware. Our advantage is physical. We designed the machines and the stack that drives them for efficiency per token, not for benchmark headlines.
- We charge less. Same open models available everywhere, at a lower rate. That is the whole pitch, and it is measurable against any other provider serving the same weights.
- We keep nothing. Prompts and completions are processed in memory. Not stored, not logged, never used for training.
Compatibility
Nothing to rewrite
The API follows the OpenAI chat completions schema. Point your existing client at our base URL and keep the rest of your code. Streaming is supported, and token usage is returned on every response so your accounting reconciles against ours.
Infrastructure
Where inference runs
Capacity is distributed across six regions, declared per endpoint so requests carrying data-residency requirements route correctly.
Company
Early, and saying so
Tannit AI LLC is registered in Haifa, Israel. We are early — a small team, a short model list, and a deliberate decision to do one thing well before doing more. If you need dedicated capacity or a region constraint, talk to us and you will reach a person.