
Cerebras unveils CS-4 server to accelerate AI chatbot responses
Cerebras Systems launches CS-4, a server rack with three large chips, to speed AI chatbot inference and ease deployment.
Cerebras Systems has introduced a new server rack, the CS-4, designed to accelerate the inference phase of AI chatbots—the process that generates answers in models like Anthropic's Claude. The system, announced Tuesday, is built around three of the company's large custom chips, which Cerebras says deliver improved performance over earlier hardware.
The CS-4 is based on Cerebras' Nexus server architecture, which uses pluggable modules to house the chips. The company's chips are notably large, roughly the size of a dinner plate, and this size helps avoid the slowdown and energy costs associated with moving data between smaller chips. The new system also includes a chip called the WSE-3 Turbo and upgraded networking components that speed data transfer between chips.
Cerebras claims the CS-4 is easier to deploy, with 50% fewer components than previous systems. Chief Technology Officer Sean Lie said at a media briefing in San Francisco that this simplification would accelerate data center construction. The server rack is available in the third quarter, and the chips are manufactured using TSMC's 5-nanometer process.
Looking ahead, Cerebras plans another generation of its chip and server in 2027. CEO Andrew Feldman said the company expects to deliver 600 megawatts of computing power by the end of that year, and that engineering efforts are focused on increasing data throughput. "We're going to get four times as fast between now and the end of 2027, and we're going to get 20 times more throughput," Feldman said.
Last week, Cerebras reported an adjusted loss of $6.9 million on sales of $180.1 million.