Business

Prime Intellect Launches Prime Inference Platform

Prime Intellect has launched Prime Inference, a platform processing 600 billion tokens daily to provide developers with highly optimized, low-latency serving for frontier open-weight models.

AlphaSignal1 day agoBusiness
Image: AlphaSignal

Prime Intellect has officially launched Prime Inference, a production-grade serving stack that currently processes nearly 600 billion tokens daily across multiple datacenters. Available via an OpenAI-compatible endpoint at api.pinference.ai/api/v1, the platform offers both serverless endpoints and reserved capacity. Its initial public deployment features the GLM-5.3 model running on NVIDIA Blackwell GB200 NVL72 systems, delivering 101 tokens per second per user across 66 concurrent sessions.

To achieve these speeds, the platform separates prompt prefill from token decoding using NVIDIA Dynamo, vLLM, Mooncake, and FlashInfer. This disaggregation reduces 90th-percentile inter-token latency by nearly 40% using a ratio of one prefill group to four decode groups. Additionally, Prime Intellect implemented NVFP4 compression for GLM-5.3's multi-head latent attention cache, shrinking cache rows from 576 to 352 bytes. This 4-bit format expands decoder capacity by 50%, raising the limit from 1.09 million to 1.63 million cached tokens.

The company also addressed latency bottlenecks in multi-node NVLink setups, which initially added 292 milliseconds to time-to-first-token compared to InfiniBand. By adopting the vLLM community's BLHNC block-major key-value layout, Prime Intellect reduced transfer descriptors ten-fold from 19,559 to roughly 1,940, cutting mean transfer times by 47% from 146 to 78 milliseconds. On prefill GPUs, switching to a DEP8 layout instead of TEP8 boosted usable prefix-cache capacity five-fold. Furthermore, halving the prefill budget from 8,000 to 4,000 tokens per GPU slashed median queue times from 550 to 110 milliseconds, yielding a 20% reduction in median time-to-first-token.

For AI practitioners, these optimizations translate to highly reliable, cost-effective agent deployments. Prime Intellect integrated a structural-tag builder into Dynamo to let vLLM's xgrammar engine enforce tool-call schemas, eliminating silent failures where models ignore tools. The team also contributed its native NVFP4 sparse-attention kernel to FlashInfer, which cuts a 35-token attention launch from 41 to 20.6 microseconds. These advancements allow developers to run long-context coding and research agents without suffering from prohibitive latency or structural errors.

This is our own summary of reporting by AlphaSignal

More in Business