Cerebras Runs Alibaba's Qwen 3.8 27B at 1,850 Tokens per Second
Cerebras is now serving Alibaba's Qwen 3.8 27B model at 1,850 tokens per second, giving developers ultra-low latency for complex reasoning and vision tasks.

Cerebras Systems has integrated Alibaba's dense, open-weight Qwen 3.8 27B model into its wafer-scale inference platform. Operating on a shared endpoint, the model achieves a throughput of approximately 1,850 tokens per second. This speed allows a 2,000-token output to generate in just 1.1 seconds, compared to the 40 seconds typical of standard 50-token-per-second platforms. This massive reduction in latency is designed to streamline complex agent workflows, tool calling, and research tasks.
The multimodal model, which processes both text and images, scored a 34 on the Artificial Analysis Intelligence Index, placing its capabilities near rival models like Claude Sonnet 4.6, DeepSeek V4 Pro, and GPT-5.6 Luna. Cerebras is offering the model at $0.99 per million input tokens and $1.49 per million output tokens. On the free trial, users get a 64K context window, a 32K maximum output, and a limit of two images per request. Paid tiers expand these limits to a 128K context window, a 40K maximum output, and up to 10 images per request.
Practitioners can leverage advanced features such as parallel tool calling, structured JSON outputs, and prompt caching. By default, the model utilizes high reasoning, which adds internal reasoning tokens to improve accuracy but increases latency; developers can disable this by setting the reasoning effort parameter to none. Because Qwen 3.8 27B uses a dense architecture rather than a mixture-of-experts design like the larger Qwen3-235B, it provides a highly uniform resource profile per token, making latency and capacity requirements easier to estimate.
This performance is enabled by Cerebras's wafer-scale processor, which avoids the communication overhead of traditional GPU clusters by keeping memory and computation on a single massive chip. However, developers should note that the shared tier is intended for experimentation rather than production. It lacks a formal service-level agreement and imposes a cap of 150,000 tokens per minute. Since a continuous stream at 1,850 tokens per second generates 111,000 tokens in a single minute, concurrent or long-running agent tasks will require upgrading to dedicated or enterprise tiers.
This is our own summary of reporting by AlphaSignal



