Notable other Hetzner

Hetzner Experiments with LLM Inference Using Qwen Model

Published
Jul 24, 2026 — 09:24 UTC

Hetzner is conducting experiments with LLM inference utilizing the Qwen model, which has 35 billion parameters and a context window size of 262K. The tests, reported by Jonas Scholz, began on July 23, 2026, and aim to assess user interaction with the model. The Qwen model operates with 3 billion active parameters and has a median time to first token of 153 ms across seven short requests. It can generate 224 output tokens per second when capped at 512 tokens. Hetzner is leveraging NVIDIA RTX 4000 SFF Ada Generation and NVIDIA RTX PRO 6000 Blackwell Max-Q GPUs, which have 20 GB and 96 GB of VRAM, respectively. Scholz noted that this is an early-stage experiment, and while Hetzner’s current offerings may not position it as a significant inference provider, the right combination of models could enhance its competitiveness. The enable_thinking option in the Qwen model is highlighted as a noteworthy feature. This follows previous coverage of AI industry developments, including OpenAI’s partnerships and the emergence of new models like GLM-5 with 754 billion parameters. For more details, visit Hacker News (AI filtered).

Turing Wire

By Callan Zhang · Jul 24, 2026 · Editorial standards →

Summarised from the primary source with AI assistance under human editorial oversight. Turing Wire is not a primary source — read the original for the authoritative account.

Source: Hacker News (AI filtered)