Hetzner is working on LLM Inference - AllTheNews.today

Hetzner is working on LLM Inference

Hetzner is experimenting with an early-stage LLM inference service that offers an OpenAI-compatible API running on its own infrastructure, currently featuring the Qwen 3.6-35B model. The company explicitly states this is not a production-ready product but rather an experiment to test user demand, system scalability, and performance without billing, SLAs, or production guarantees. Initial tests show fast performance with 153ms median time-to-first-token and 224 output tokens per second, though these metrics are preliminary and not representative of real-world multi-user scenarios.
Read Full Article →
sliplane.io
← Back to Latest