- You're building an AI product. Inference is in your hot path, not a side feature.
- You're tired of paying one vendor for compute, another for Postgres, and a third for tokens — all on separate bills.
- You'd use an OpenAI-compatible API for open-weight models like Llama, Qwen, or Mistral at a meaningfully lower cost than closed frontier models.
- You want a managed Postgres, MySQL, MongoDB, or Redis instance sitting next to your app, not three hops away.
- You want to fine-tune an open-weight model on your own data and get it behind an endpoint without renting a GPU box yourself.
One account.
Every layer of an AI app.
Most clouds are good at one thing. Ours runs your app, your database, and your model inference side by side, under one identity and one bill — so stitching three vendors together isn't a tax you pay before you ship.
When NevTan Cloud helps. When it doesn't.
- Your app doesn't use AI and isn't going to. A pure-frontend deploy may be simpler on a frontend-only host.
- You need a hyperscaler contract with FedRAMP or specific marketplace SKUs. We're not procurement's first answer today.
- You run bare-metal Kubernetes and like it. Our abstractions will feel limiting.
- You only need closed frontier models (GPT-4o, Claude Opus, Gemini). We serve open-weight models — frontier model passthrough isn't live.
- You require on-prem deployment today. We're cloud-first.
Not a roadmap. Shipped, and running today.
App Platform
Git-based deploys by branch or commit, Docker builds, custom domains, automatic SSL and routing.
Managed Databases
PostgreSQL, MySQL, MongoDB, and Redis. Three sizes each, from a 1 vCPU starter to a high-memory tier with point-in-time restore.
AI Inference
An OpenAI-compatible endpoint in front of open-weight chat, embedding, and audio models, running on your own deployed worker.
GPU Marketplace
We provision and size the GPU for you — spin up instances to run model servers or training jobs, no separate GPU host to manage.
Fine-tuning
Bring a dataset, fine-tune an open-weight model, and deploy the result behind its own inference endpoint.
AI Agents
One-click hosting for the Hermes Agent — deploy, configure channels, and monitor it like any other workload.
RAG collections
Vector collections backed by MongoDB with OpenAI-generated embeddings, for retrieval-augmented generation.
Sandboxes
Disposable browser-based sandbox environments for running and testing code.
Team & security
Role-based access, audit logs, two-factor auth, and GitHub / Google / Bitbucket sign-in.
One endpoint. Real models.
Priced per token.
NevTan Inference is an OpenAI-compatible API in front of a GPU worker you deploy — not a shared black-box endpoint. You pick the model, we provision and size the GPU for you, and your app calls it the same way it would call OpenAI: change the base URL and the key.
- Llama 3 8B · $0.20/$0.60 per M
- Llama 3 70B · $0.90/$0.90 per M
- Mistral 7B v0.2 · $0.25/$0.75 per M
- Mixtral 8x7B · $0.60/$0.60 per M
- Qwen 2.5 72B · $0.90/$0.90 per M
- DeepSeek R1 Distill 7B · $0.30/$0.60 per M
- BGE-M3 · $0.02 per M tokens
- Whisper Large v3 · $0.006/min
- Fine-tune and deploy your own checkpoint behind the same endpoint pattern
→ Prices per 1M input/output tokens where both apply. Full catalog and live pricing in the dashboard.
Still on the fence?
Try it before you switch.
Deploy a side project on the free tier, add a payment method to unlock your signup credit, and call a few thousand tokens of an open-weight model. See how it feels.