Blog

Restate raises $20M Series A to make Durable Execution a building block for every backend

September 30, 2026

Stephan Ewen

We raised a $20M Series A led by Singular with participation from Redpoint Ventures, and Capital One Ventures to make Durable Execution a widespread building block for backends, rather than something confined to expensive, high-overhead workflow runtimes.

We started Restate over three years ago over our frustration about how hard it is to build any non-trivial piece of backend logic. Control planes, workflows, API orchestration. We (and everyone else) were solving the same problems in the application logic again and again. Too many times we accidentally built a poor-man's workflow engine into an application, a distributed state machine or an actor system on top of DBs, queues, locks.

It felt like a fundamental infrastructure primitive was missing.

When we built Apache Flink, one of the core ideas was to make stateful stream transformations a building block from which you could compose a wide range of streaming applications. We wanted something similar for the asynchronous, event-driven, long-running work that increasingly makes up our backends.

Fast-forward three years, and that problem has only become more important. AI agents made long-running asynchronous programs mainstream. A single agent can perform hundreds of model calls, tool calls, waits, retries, callbacks, and require human approvals. The same problems that you have in workflows and backend orchestration now show up inside agent loops, often at much higher frequency.

Durable Execution has emerged as one of the answers to this problem: It lets you write application logic while the runtime takes responsibility for recording progress, recovering from failures, retrying work, and recovering partial progress.

But for all the progress the category has made, we think it still falls far short of its potential.

Durable Execution is the starting point, not the destination

The original generation of Durable Execution systems was designed primarily around workflows. That design comes with assumptions about latency, cost, deployment, and programming model, and becomes limiting when you try to use durability throughout a backend. We think four things need to change:

First, Durable Execution needs to become more than a workflow tool.

Workflows are important, no question about that. But look at everything happening inside a modern backend and much of it does not naturally look like a workflow. You have services talking to services, stateful entities, coordination around shared state, sessions, RPC, events, queues. These things need durability too, in their communication, state management, concurrency control, flow control. That isn’t captured by workflows and activities.

One example of this is Restate’s Virtual Objects, which have quickly emerged as a fan-favorite among Restate users for anything from agent context to session stores, or custom queues. They make it easy to build stateful concurrent entities without having to jump through hoops and program with signals and continue_as_new.

Three chat Virtual Objects with independent session state, beside their shared TypeScript definition.

A Virtual Object is a stateful entity keyed by an ID, for example a chat session, with durable state and one write handler running at a time per key.

That is where we think the category needs to go.

Second, Durable Execution needs to become much more efficient.

Managed Durable Execution is commonly priced in the tens of dollars per million durable actions (often $25 - $50). In comparison, databases and queues commonly cost around ~$1 for million writes or events.

Of course, a durable action does more than a database write or a queued event, but this order of magnitude difference still matters, because at high per-action prices, developers naturally use durability sparingly and create coarse activity boundaries.

When durable execution becomes more efficient, applications use more of them and get simpler semantics and better debugging capabilities: durability inside the agent loop (persisting inference, tool calls, guardrail evaluations) ensures you always get a consistent execution history, compared to durability around the loop; that matters for agents performing sensitive, long-running, or expensive tasks.

On failure, durability around the loop restarts the activity from the beginning; durability inside the loop reuses saved results and retries the interrupted step.

We see this already happening with Restate, users start to build programs that use durability more generously.

Third, Durable Execution needs much lower latency.

This goes hand-in-hand with the second point. Durable actions incur not just monetary cost (directly, or through their resources requirements), but also a cost in the latency-budget of the logic. Getting durable actions from 25+ms to less than 5ms makes the decision to introduce an extra step easier.

Fourth, Durable Execution needs to be much lighter to operate and integrate.

Durable execution runtimes that are scalable and highly available consist typically of multiple systems plus a database and are complex to run (or nudge you to their SaaS offering).

Restate with Kubernetes, Cloud Run, Cloudflare Workers, and other deployment targets compared with a traditional durable execution cluster containing Frontend, Matching, History, and internal Worker services, a separate database cluster, and separate application worker instances.

That is a problem for sensitive applications (e.g., financial- and healthcare workflows, agents handling loan origination or medical bill disputes) that need to keep the data within a network boundary. And ironically, these sensitive apps are among those that benefit most from the rigor and predictability of running on durable execution.

That complexity is not inherent to durable execution, but an artifact of how systems were built. Restate’s design (dependency-less binary that forms clusters) shows an alternative way of doing it, which opens up viable self-hosting and BYOC deployment options.

What happens when durable execution becomes fast, efficient, beyond workflows?

Over the last year, we have seen customers start building systems on Restate that changed our own understanding of how broadly Durable Execution can be used.

Replit

Replit moved the execution of the Replit Agent onto Restate as part of a larger evolution of the Agent architecture. Their new architecture uses more than 10x more durable actions as the previous generation.

That only works if a durable action is fast and resource-efficient enough to become part of the inner agent loop rather than a coarse workflow boundary. Instead of bouncing between a workflow orchestrator and separate activities, an agent can stay inside a process while individual tool calls, inference calls, and other operations become durable. The process gets durability without giving up the programming model that makes a tight agent loop useful.

DOSS

DOSS runs high volumes of enterprise workflows and subworkflows on Restate. It replaced their BullMQ-based architecture, improving latency, failure handling, observability, and the ability to understand and debug executions.

Fortune 500 bank

A Fortune 500 bank adopted Restate for financial workflows with multi-region consistency requirements, a very different workload with very different constraints.

These systems have little in common at the application level. But what they share is that durability is not a special feature applied coarsely, but a basic execution model of the application.

That is what we are doing and where we think the space is heading.

Breaking a few rules

Building Restate required us to break a few rules that have become conventional wisdom in infrastructure.

Rule #1: Don't build your own storage system

We half-broke this one. Restate uses RocksDB for local storage and object stores for durable snapshots. But around that, we built the replicated log, consensus and failover handling, a query engine, and a processing model specifically for Durable Execution.

The reason is that Durable Execution has a very specific access pattern: On the hot path (make a function/step durable), it is actually very simple (don’t let the VCs read this): you only need to be able to durably append to a log really fast. In Restate, that is optimized to boil down to a quorum write (a single network roundtrip), which is the minimum work you can possibly do, while preserving full durability. Data access is served from RocksDB, which materializes the orchestration state by following the log.

Building that was a big decision. We could have assembled Restate on top of an existing database, queue, and consensus-backed metadata system. It would certainly have been easier to get started. But the bespoke design gave us a major boost in efficiency, latency, and availability. Plus, the operational benefit of fewer moving parts.

Rule #2: Task queues and polling workers are the only way to build scalable orchestration

Workflow engines usually put work into task queues and have workers pull from them. Restate inverts that model by pushing invocations to durable function deployments (containers or FaaS).

The push model is much harder to get right, because it doesn’t get the inherent back-pressure that workers have by pulling work from queues slower. So you need to build sophisticated and efficient flow control into the dispatcher (which we did!)

To make the most of it, you want the connections to be both fast/streaming (for low latency steps) but also support the code-that-sleeps-for-a-day pattern. To solve that, Restate uses a fast bidirectional stream and implements a suspension protocol for the durable functions. Finally, Restate Cloud also allows you to let services connect into the cloud, or deploy a local dispatcher to avoid exposing public endpoints.

All that is harder to build and get right, but once it works, we believe it is much better: It supports serverless functions (as well as containers), it supports really low latency steps, it gives you the flexibility to run your work in a workflow pattern (workflow / activity separation) or as durable processes (the later being a great match for AI agents)

And feels like you are deploying functions, which is where the infra world is going. Try integrating Restate with CloudFlare, Vercel, Lambda, CloudRun vs. a workflow-based runtime. You will see a big difference.

Rule #3: State outside workflows belongs in databases

That is mostly true. But there are many use cases where you naturally need stateful concurrent entities that are interacted with from durable functions: agents, sessions, explicit state machines, deployments, digital twins. Basically anything where state lives beyond a single workflow, but is interacted with mostly from workflows.

The Virtual Objects do that in Restate and allow you to build stateful logic elegantly. They are deeply integrated into the execution of durable functions to give you seamless exactly-once semantics for state mutations. Deploy them on FaaS, and you get stateful-serverless functions with a simple concurrency/consistency model that work as an amazing building block for many applications.

The space is much bigger than Durable Execution

We believe something like Durable Execution will become a fundamental part of how backends are built. But for that to happen, it cannot remain an expensive tool that you bring in for a few particularly complicated workflows.

It needs to become efficient enough to use pervasively, fast enough to put into latency-sensitive paths, simple enough to deploy wherever your applications run, and flexible enough to model more than workflows.

Restate has made big progress in that direction. Its abstraction with Durable Functions, Virtual Objects, Virtual Queues, durable RPC feels too useful for almost any piece of async work, anything that would traditionally run behind a queue, or calls into multiple services. And its distributed log-based architecture reduces latency, cost, and operational complexity.

Long term, we think the interesting opportunity is bigger than building a better workflow engine. We want to build the runtime underneath agents, workflows, and distributed applications that makes durability, state, communication, and coordination feel like normal programming primitives.

The Series A gives us the resources to push that idea much further.

There is still a lot left to build.

Keep reading