Scale up instantly, scale down effortlessly.
ALPCRUN.AI is the AI inference engine for your business. We provide dedicated GPU clusters and host enterprise-grade models for companies that need compliant, high-performance AI inference in Europe. From open-weight models to custom fine-tunes, we provide the full runtime environment – scalable on demand, billed by GPU-time and set up in days, not quarters. Connect your apps to a secure, compliant European inference backend and start scaling, exactly as your workflow needs.
Your AI runs on someone else’s terms.
Today that usually means a US cloud, a shared API with somebody else’s queue, and a GPU bill that keeps running while nothing runs. Serious inference deserves a backend that scales with your traffic, lives where your data is allowed to live, and does not need an ops team to keep it up.
Capacity that follows your traffic
Dedicated GPUs spin up when requests arrive and scale back to zero when they stop. You pay for GPU-time that served a request, not for hardware waiting for one.
Compliant by location, not by promise
Frankfurt and other EU regions, GPUs reserved for your company alone, and nothing used to train anyone’s model. The question from your data protection officer has a one-sentence answer.
Live in days, no ops team
Pick a region and a model, or upload your own weights. We size the hardware, run the serving stack and keep it current. Your developers change one base URL.
Everything an inference backend should do.
The unglamorous parts of running models in production, handled for you: hardware, serving, scaling, residency, and a bill that reads the way the month ran.
OpenAI-compatible endpoint
Chat, embeddings, streaming and tool calls behind the API your code already speaks. Change one base URL, keep your SDKs and frameworks.
Dedicated GPU clusters
Per-customer clusters with their own queues and KV-cache. One GPU up to multi-node for the largest open models, created in minutes.
Scales with load, down to zero
Replicas follow your traffic within the limits you set. An idle cluster scales to zero and stops billing, then wakes when the next request lands.
Open-weight models, kept current
Mistral, Llama, Qwen, Gemma, Kimi, DeepSeek and friends, sized and tuned by us. Swap a model without changing anything in your applications.
Your own fine-tunes
Upload your weights to a private, encrypted registry. They are pulled only by your own cluster and served next to the public models.
Hosted in the EU
Requests are served in Frankfurt or the EU region you choose, by a German company, and never used for training.
Billed by GPU-time
Per GPU-hour on dedicated clusters, per token on shared endpoints, itemised per cluster and key. No idle GPUs on the invoice.
Want to see it running on your models?
Register for early access →Start shared. Scale dedicated.
Most workloads should not start on reserved GPUs. Start on shared European endpoints and pay per token for as long as that serves you. Move a model to hardware of your own when latency, volume or a regulator makes it necessary. Both run in the region you choose, Europe by default, behind the same endpoint.
Shared inference endpoints
Nothing to provision. Your application works the same afternoon and you pay for the tokens it actually used.
Your own private cluster
Dedicated GPUs in the region you choose, for models a shared endpoint cannot keep up with or a regulator will scrutinise.
From sign-up to first token.
Three steps, no ops hires, no procurement cycle for hardware. You need one person who can choose a model and one who can approve a GPU-hour budget. They can be the same person.
Choose the setup
Decide where it runs, what it serves, and what it may cost.
Connect your applications
One base URL and a key per application or environment. Your code does not change.
Watch it scale
Traffic arrives, replicas follow, and the invoice shows exactly what ran.
An answer to “what does inference cost us?”
With reserved GPUs the honest answer is usually “the same, whether we use them or not.” Here the invoice reads the way the month ran: GPU-hours that served requests, tokens on shared endpoints, itemised per cluster and key.
# billed for what ran, itemised per cluster
cluster model peak gpu-hours idle
prod-chat Llama 3.3 70B 12 1,840 0
legal-v3 your fine-tune 4 310 0
staging Mistral Small 3.2 2 96 0
shared per token · 38.2M tok – – –
total dedicated gpu-hours 2,246 0
- ▸Priced the way it runs: per token on shared endpoints, per GPU-hour on your own cluster. No reserved instances you have to keep busy.
- ▸Zero when idle: a cluster that scales to zero bills a low flat baseline and nothing else until the next request.
- ▸Caps before surprises: a maximum GPU count per cluster and a monthly ceiling per key or environment mean the invoice never arrives as news.
- ▸Itemised, not averaged: GPU-hours and tokens per cluster, per key, per environment, ready for whichever cost centre they belong to.
- ▸Euros, one invoice: one European supplier, one invoice in euros, one data processing agreement.
Your data does not leave Europe unless you decide it does.
Not a checkbox somebody has to remember to tick. It is simply where the hardware is. If a team ever needs a US frontier model, that is a deliberate step, taken in Torhaus.AI, for one team and one job.
- ▸Served inside the EU: shared endpoints in Frankfurt, or GPUs reserved for your company alone. Either way the request does not leave.
- ▸Never used for training: your prompts and documents answer your questions and do nothing else.
- ▸GDPR without the essay: one region, one processor, and a data processing agreement you can actually sign.
- ▸Audit-ready by construction: every response carries the region, cluster and billing mode it ran on, and every request is logged as one row you can hand to an auditor.
- ▸Frontier models are a separate door: GPT and Claude are not served here. If a team needs them, Torhaus.AI attaches them per team, labelled as leaving Europe.
- ▸German and English: the dashboard and the documentation come in both, and your invoice reads in euros.
Models are a menu, not a strategy.
Nobody should have to follow model releases to run a product. You pick the job, we keep the menu current and the serving stack tuned, and you decide which models your applications may call.
- ▸Open models, either tier: Mistral, Llama, Qwen, Kimi, Gemma, DeepSeek, and friends. A small model on a single GPU, or one of the biggest that needs multiple racks of them. Sizing the hardware is our job, not yours.
- ▸Your own fine-tunes: upload your weights into the private registry we host and serve them next to the rest.
- ▸Frontier models, via Torhaus.AI: GPT and Claude are not on this menu. Torhaus.AI attaches them per team next to your ALPCRUN.CH models, labelled as leaving Europe.
- ▸Sensible defaults: a well-sized everyday model handles most work, so nobody compares benchmarks to ship a feature.
Need the governance layer? That is Torhaus.AI.
Torhaus.AI grew out of ALPCRUN.CH AI and is now a product of its own: an AI gateway that gives every person and team a key, a budget and a set of rules, in front of whichever models your company uses. ALPCRUN.CH AI is one of the backends it points at. Together they cover the whole path from an employee’s prompt to a GPU in Europe.
- ▸Point it at us in one setting: Torhaus.AI adds an ALPCRUN.CH cluster or shared endpoint as a backend the way it adds any OpenAI-compatible provider. Nothing changes for your users.
- ▸Governance lives in Torhaus.AI: virtual keys per person and team, budgets and caps, PII guardrails, and audit reports mapped to the EU AI Act. None of it slows the backend down.
- ▸Frontier models when a team needs them: GPT and Claude attach in Torhaus.AI per team, always labelled as leaving Europe. Everything on ALPCRUN.CH stays in the EU.
- ▸Use either on its own: ALPCRUN.CH AI serves any OpenAI-compatible client, and Torhaus.AI routes to any provider. You are never locked into the pair.
- ▸Built by the same team: both come from Lime Labs, so they are tested together and the metering on both sides agrees.
The details, for the people who ask for them.
Under the friendly surface is a real inference platform: dedicated GPU clusters running hosted llm-d, shared endpoints for the everyday, and one OpenAI-compatible API in front, metered and priced in one place. If you want governance on top, Torhaus.AI sits in front of it.
Private dedicated clusters
- Per-customer GPU clusters, never shared queues
- Created and torn down in minutes
- One GPU up to multi-node for the largest open models
- Scale to zero when idle
- Your models, your KV-cache, your GPUs
Smart request routing
- KV-cache-aware and load-aware steering
- Disaggregated prefill and decode pools
- Multi-region clusters for failover
- More throughput from the same hardware
Keys, limits and metering
- One key per application or environment
- Per-key model allow-lists, RPM and TPM limits, expiry
- Usage per key, cluster and model, live
- Your clusters, your keys, never a shared control plane
One row per request, always
- Identity, cluster, model, tokens and cost
- Region and billing mode, echoed in the response headers
- Per-stage latency for capacity planning
- Message bodies only when you enable them
- Written asynchronously, never on the request path
OpenAI-compatible API
- Change one base URL, keep your code
- Chat, embeddings, images, audio, and moderations
- Chat frontends and coding agents, unchanged
- Streaming and tool calls included
- Works with existing SDKs and frameworks
Metered on what you run
- Per token on shared endpoints, per GPU-hour on your own
- Itemized per cluster, environment, and key
- Maximum GPU count and spend caps per cluster
- Low flat baseline when scaled to zero
Private model registry
- Upload your own weights from the dashboard
- Versioned and encrypted, hosted by us
- Pulled only by your own cluster
- Runs beside public models in one pool
Torhaus.AI in front, if you want it
- Keys per person and team, budgets and caps
- PII guardrails, deny-lists, block or redact
- Audit reports mapped to EU AI Act and DSGVO
- Adds ALPCRUN.CH as a backend in one setting
Leaving costs you a base URL
- The OpenAI API you already write against
- Open-weight models, on llm-d and vLLM, that run anywhere
- Standard request and response shapes, no proprietary SDK
- Your fine-tuned weights stay yours
# OpenAI-compatible. Any model. Your own GPUs, in Europe.
$ curl https://inference.alpcrun.ch/v1/chat/completions \
-H "Authorization: Bearer $ALPCRUN_KEY" \
-d '{ "model": "mistralai/Mistral-Small-3.2-24B",
"messages": [{ "role": "user", "content": "…" }] }'
< x-alpcrun-region: eu-central (Germany)
< x-alpcrun-cluster: c-9f2a (private, dedicated)
< x-alpcrun-replicas: 12 (autoscaled, max 16)
< x-alpcrun-billing: gpu-time (no token charge)
Put your inference in Europe.
Register for early access and we set it up with you: region chosen, models deployed or your weights uploaded, endpoint live. Usually a few days, and you know the GPU-hour rate before you commit.
Going to Bits & Pretzels 2026 in Munich? Come by the booth for the two-minute demo.