Independent , Honest and Dignified Journalism

Cerebras to Supply AI Systems to Gimlet Labs as Demand for Fast AI Inference Expands

The partnership will provide Gimlet Labs with Cerebras computing systems designed to deliver high speed AI inference for applications including cybersecurity, voice technology and financial analysis.

SAN FRANCISCO, Sept 29: Cerebras Systems will supply artificial intelligence computing systems to cloud startup Gimlet Labs under a new partnership aimed at expanding access to high-performance infrastructure for AI applications.

The agreement, announced on September 28, involves Cerebras providing AI chips and related hardware with capacity of roughly 100 megawatts to Gimlet Labs. The startup plans to use the systems as part of its cloud infrastructure and make the computing capacity available to customers developing products based heavily on artificial intelligence.

The deal highlights the growing importance of AI inference, the stage at which trained models process new information and generate responses. While much of the attention around AI infrastructure has focused on training increasingly large models, companies are now also seeking faster and more efficient systems for running those models once they are deployed.

Gimlet Labs plans to use Cerebras systems to support frontier AI models that require substantial computing resources. The startup expects the hardware to become available through its cloud platform in 2027, while Cerebras is expected to deliver the systems over a period of one to two years.

Cerebras CEO Andrew Feldman said the arrangement demonstrates the ability to deploy the company’s systems in environments containing hardware and infrastructure from different suppliers. Gimlet CEO Zain Asgar said rapid inference could be particularly useful for businesses developing applications in areas such as cybersecurity, voice technology and financial analysis.

The companies did not disclose the financial terms of the agreement.

The partnership comes as the technology industry moves toward AI systems that can generate responses and complete tasks in real time. Inference speed is becoming increasingly important because users expect AI-powered applications to respond quickly, even when the underlying models are large and computationally demanding.

For cloud providers, faster inference can affect the economics and usability of AI services. A system that can produce responses quickly can support applications where delays may affect the user experience or the usefulness of the service.

Cybersecurity is one area where rapid processing can be particularly significant. AI systems can be used to examine large quantities of information, identify patterns and assist security teams in analysing potential threats. The ability to process information quickly can help applications respond to changing circumstances.

Voice-based applications also require rapid inference. Delays between a user’s speech and an AI system’s response can make interactions feel unnatural. As developers use larger models for voice assistants and other conversational products, the underlying computing infrastructure becomes an important part of the experience.

Financial applications represent another potential use case. AI systems can process large amounts of market, transaction and business information, while financial services companies are increasingly exploring automated analysis and decision-support tools.

Gimlet’s strategy is to offer infrastructure to companies that are building products around AI rather than treating artificial intelligence as a secondary feature. Its customers are expected to include businesses and startups that require substantial computing capacity but may not want to build and operate their own large-scale AI infrastructure.

The partnership also illustrates the increasing competition among specialised AI hardware providers. Nvidia remains a dominant supplier of computing infrastructure for AI workloads, while companies such as Cerebras are developing alternative architectures aimed at improving performance for particular types of workloads.

Cerebras has focused heavily on the speed of AI inference. Its systems are designed around large-scale processing capabilities intended to reduce the time required to generate outputs from advanced models.

The company’s CS-4 systems, which are part of the new agreement, were unveiled earlier in 2026. Cerebras plans to provide the systems to Gimlet over the next one to two years, with Gimlet targeting 2027 for their integration into its cloud services.

The arrangement is also significant because AI infrastructure is becoming increasingly heterogeneous. Cloud companies may operate hardware from multiple vendors, with different processors, networking systems and software environments working together.

For customers, that can create both opportunities and technical challenges. A broader range of hardware can allow cloud providers to match specific workloads with suitable systems, but operating different architectures can require additional software and infrastructure management.

Cerebras said its technology can be deployed alongside systems from other manufacturers. The company sees this flexibility as an important feature as cloud operators build infrastructure around several types of AI processors rather than relying on a single supplier.

The growing focus on inference is also linked to the changing nature of AI applications. Earlier generations of AI infrastructure were heavily shaped by the need to train models using enormous datasets. Once a model has been trained, however, it still requires substantial computing power every time it generates an answer.

As millions of users and businesses begin interacting with AI systems, those repeated inference workloads can become a major component of overall computing demand.

This has encouraged technology companies to develop hardware and software specifically aimed at making inference faster and more efficient. Cloud providers are responding by expanding specialised infrastructure that can support customers without requiring them to build their own data centres.

The Cerebras-Gimlet agreement forms part of this wider infrastructure race. Gimlet intends to use the new capacity to serve companies developing AI-focused products, while Cerebras gains another deployment for its specialised systems.

The two companies did not provide a detailed timetable for each stage of the installation, beyond saying that Cerebras would supply the systems over one to two years. Gimlet expects to make the hardware available through its cloud platform in 2027.

The scale of the planned deployment also illustrates the energy requirements associated with modern AI computing. The systems covered by the agreement are expected to have capacity to consume approximately 100 megawatts, underlining the amount of electricity required by large computing installations.

Energy use has become an increasingly important issue for the technology sector as companies expand AI data centres around the world. Greater computing capacity can enable more sophisticated models and faster responses, but it also increases requirements for electricity, cooling and physical infrastructure.

For cloud operators, the challenge is therefore not simply acquiring powerful processors. They must also build the supporting infrastructure needed to operate those systems reliably.

Gimlet’s planned deployment shows how AI infrastructure providers are positioning themselves for a market in which companies want access to advanced models without necessarily owning the hardware on which those models run.

The partnership also points to a growing separation between AI model development and AI computing services. Model developers can concentrate on creating software, while specialised cloud providers supply the computing resources required to operate those systems at scale.

As AI applications move into areas such as security, finance, communications and enterprise software, response speed is likely to remain an important factor in how these products are designed.

The Cerebras-Gimlet deal consequently represents more than a hardware supply agreement. It reflects the expanding infrastructure layer supporting the next phase of AI adoption, where the ability to run powerful models quickly and at scale is becoming as important as developing the models themselves.

WhatsApp Channel