When a consumer asks Ask DoorDash to “reorder my usuals,” the request sounds simple. DoorDash already has their order history; why not retrieve the products they buy most often and put them into a shopping list?
Reordering is one of our most popular use cases, accounting for roughly one in every three orders. But the original agentic path could take tens of seconds to produce a list. By moving stable historical reasoning ahead of the request and using deterministic execution where the product already knows the consumer’s intent, we reduced one reorder path from roughly 16 seconds to roughly 2 seconds and doubled the rate at which consumers viewed the resulting list.
Moving the historical reasoning offline also relaxed the latency constraint on generation, allowing us to use more capable models to produce higher-quality usuals. That improvement translated into a ~34% increase in order rate for the flow.
The challenge is that making reorder faster cannot mean simply replaying history. Order history records what someone bought, not what they need next.
Milk purchased once a week is probably a recurring need. A large pack of paper towels purchased at roughly the same frequency may not be. Six units bought in one delivery are different evidence from one unit bought across six deliveries. A yogurt a consumer buys repeatedly may be a staple; three flavors bought once each may be experimentation.
Answering “shop my usuals” therefore requires judgment over historical behavior. At the same time, the final answer has to be grounded in a catalog that changes continuously: items go out of stock, listings change, and the exact product a consumer bought previously may not be available at the store they are shopping today.
For a typical free-form reorder request, the Grocery agent loads the Reorder skill, retrieves the consumer context and order history needed for the request, decides which historical purchases belong in the next basket, resolves those needs against the current catalog, and assembles the shopping list. This flexibility lets the same flow handle a range of reorder requests while bringing together several different parts of the Ask DoorDash system to produce a seemingly simple shopping list.
Earlier posts in this series covered Intelligence, Evaluation, Platform, and User Experience one at a time. Reorder is the flow where all four have to agree at once: the list has to be personal, verifiable, fast, and shoppable. This post follows a single request through each of them.

Reorder is more than replay
Ask DoorDash supports a range of reorder requests. A consumer might type “What groceries do I usually get?”, “Reorder my usual groceries”, “Which milk did I buy last time?”, or “Reorder my usuals, but cheaper”.
These requests sound similar, but they require different decisions underneath. Looking up a product from the last order is different from deciding which purchases represent a consumer’s usuals; asking for cheaper versions adds another set of constraints.
Ask DoorDash’s multi-agent architecture, described in our engineering overview, first routes a grocery request to the Grocery agent. When the request involves reorder, the agent can load the Reorder skill.
The Reorder skill gives the agent task-specific instructions for handling these requests using the context and tools available to it. Rather than defining a separate workflow for every possible phrasing, it gives the agent a general process for deciding what the request requires and turning the result into an actionable shopping list.
Conceptually, that process has five stages:
- Establish context. Determine the store to shop at when one is not already known and retrieve relevant context about the consumer.
- Retrieve history. Read the grocery order history needed for the request and reshape it into a representation the model can reason over.
- Select what matters. Decide which parts of that history answer the consumer’s request. For “usuals,” this means distinguishing recurring needs from one-off purchases.
- Resolve products. Search the current DoorDash catalog, preserving details such as brand, size, and variant wherever possible.
- Build the result. Ask the consumer a clarifying question when a genuinely ambiguous choice remains, then create the editable shopping-list artifact.
Other skills can compose with that flow. If a consumer asks for their usuals “but cheaper,” for example, the Reorder skill can establish what they typically need while affordability reasoning considers the products and prices available now.
This flexibility matters for free-form language. A consumer typing into Ask has not selected an API operation; they have expressed an intent, and the agent has to determine what work that intent requires.
Every one of those five stages assumes something the skill does not contain. Which store to shop, which brand to prefer, whether this consumer even wants olive oil this week — none of that is a fact about reorder. It is a fact about the consumer.
Intelligence: What the agent already knows about you
Order history tells us what a consumer bought, but it does not tell us how to interpret every purchase. As described in Building Ask DoorDash (Part 2): Intelligence, Ask DoorDash can bring in context such as store preferences, brand affinities, dietary preferences, and pantry state from long-term, conversational, and in-session memory.
For reorder, those signals are not another source of products. They help interpret the consumer’s purchase history. If someone consistently buys one milk brand, for example, that preference should carry into the resulting list. If they recently told Ask they are well stocked on olive oil, that context can keep olive oil out of their next reorder even if it appears frequently in their history.
That matters because consumers already know what their own “usuals” are. Small mismatches in store, brand, or pantry state are immediately visible and can quickly undermine trust in the rest of the list.
Intelligence therefore gives the Reorder skill a better starting point, but it does not answer the core question: which purchases in the consumer’s history actually count as recurring needs? That decision still has to be made — and evaluated consistently.
Evaluation: Testing whether the list is right
We needed a way to evaluate the “usuals” decision before it reached the consumer. As described in Building Ask DoorDash (Part 3): Evaluation, agent evaluation looks at both the execution path and the result. For reorder, that means checking whether the agent retrieved the right context and history, followed the expected flow, and produced an actionable shopping list — as well as whether the items on that list actually reflect the consumer’s recurring needs.
Those selection checks are concrete. If a consumer repeatedly orders three yogurts, does the list preserve that pattern and quantity? Does an item that appeared only once stay out of their “usuals”? If the consumer has said they already have something at home, is that reflected in the result?
Because live order history changes over time, we use the fixture-based Conversation Simulator to replay reorder scenarios against a fixed purchase history. Checklist-style rubrics then let us test recurring items, one-off purchases, quantities, and known preferences against the same evidence as the implementation changes.
A correct list is only useful if it arrives while the consumer is still shopping. Our original flow took about 16 seconds.
Platform: The decision doesn't need to happen while you wait
That latency came partly from the way Ask DoorDash handles agentic requests. As described in Building Ask DoorDash (Part 4): A Platform for Building and Evolving Agents, within a single consumer turn, the model can make a decision, call a tool, inspect the result, and use that information to decide what to do next. In a representative Grocery flow, that can require 6–8 LLM calls along with several tool calls.
That flexibility matters when each step depends on information discovered during the request. However, reorder showed us that several of those steps relied on historical data that changes slowly.
Consumers often return to a relatively small set of grocery stores. Instead of selecting a store, fetching its order history, and determining the consumer’s recurring needs on every request, we can do that work ahead of time using historical data.
Other parts of the flow have to stay live. Inventory and prices change, products become unavailable, and catalog listings can change between visits.
That gave us a natural boundary: precompute the consumer’s likely stores and recurring needs, then use the live request path to resolve those needs against the current catalog and build the shopping list.
That split became the basis for our offline-to-online pipeline.

Figure 1: Ask DoorDash assigns each reorder request to an execution path based on where its judgment belongs. For “reorder my last cart,” a typed session scope allows a backend plugin to retrieve the previous cart without using the model for tool selection. For recurring purchases, a daily pipeline combines 90 days of order history with task-relevant profile signals, selects consumers with sufficient evidence, runs batch LLM inference, and validates the output into a reusable usuals bundle. “Reorder my usuals” reads that bundle directly, while “reorder my usuals, but cheaper” supplements it with current prices, promotions, package sizes, inventory, and consumer constraints. The three paths converge on live catalog resolution, which attempts the store item ID first, then the catalog product ID, and finally a search term before creating an editable shopping list. The Grocery agent remains available to clarify ambiguous choices, narrate substitutions, and recover from missing or unusable results.
Generating usuals before the consumer asks.
Cohort selection and input materialization
Not every order history contains enough evidence to infer recurring needs. A single grocery delivery tells us what a consumer bought once, not whether any of those items are staples. Running generation across every consumer would therefore add cost while producing low-confidence results.
We treat cohort selection as both a quality gate and a cost-control mechanism. A daily Spark job identifies consumers with enough recent grocery or retail activity to support generation.
For each selected consumer, we prepare two inputs: a compact profile containing the long-term signals most relevant to reorder, such as item-category interests and brand affinities, and item-level order history from the previous 90 days.
The two serve different purposes. The profile provides context about the consumer; the order history provides the event-level evidence needed to infer recurring needs. A brand preference may tell us which yogurt a consumer is likely to want, for example, while purchase history tells us whether yogurt belongs in their usuals at all.
Following the task-specific context trimming approach from our work on offline LLMs for online personalization, we send only the profile signals relevant to this task rather than the consumer’s full profile. That reduces input cost and keeps unrelated context from competing with the purchase evidence the decision depends on.
Defining “Your usuals”
Producing a ranked list is straightforward once “usual” is defined. Defining it was the harder part.
Frequency is a reasonable baseline, but it answers what a consumer bought most often rather than what they are likely to need again. Milk and a large pack of paper towels can appear in the same number of orders while having very different replenishment patterns. Six bottles in one delivery are also weaker evidence of a recurring habit than one bottle purchased across six separate deliveries. Recency alone does not resolve the ambiguity either.
We use an LLM because these signals interact in ways that are difficult to capture in a single ranking rule. The model follows an explicit qualitative ordering: frequency across distinct deliveries is the primary signal of a recurring need, recency indicates whether that need is still active, and replenishment characteristics help distinguish quickly consumed products from durable ones. Brand affinity can refine the choice, but it does not outweigh direct purchase evidence.
Stock-on-hand adds another layer. A recently purchased multipack of a durable product may move down the list, while a perishable that appeared regularly in the past but has not been purchased recently may move up.
Turning the decision into a bundle
The model applies that reasoning independently for each relevant store and returns the recurring items it identifies. The maximum number of items is a ceiling, not a target: if the evidence supports only a handful of usuals, the model returns the smaller set rather than filling the list with weaker guesses. If the history does not support a useful bundle, it can abstain entirely.
The output is constrained to evidence from the input. Store and item identifiers must come from the consumer’s history, repeat purchases are preferred over one-offs, and near-duplicate variants are collapsed rather than filling multiple positions in the bundle. For each item, we retain signals such as purchase frequency, recency, and typical quantity so the result can be served without redoing the historical reasoning online.
Represent the need, not the SKU
A well-ranked item is still unusable if the catalog changes before the consumer comes back. The exact listing may be relisted under a new merchant-specific identifier or disappear entirely. Persisting only the historical store item ID would make the bundle brittle; persisting only a generic term such as “oat milk” would lose the specificity of what the consumer actually buys.
We therefore represent each need at three levels: the store item ID for the most precise match, a catalog-level product identifier for the same underlying product across listings, and a generic search term as a fallback. At request time, the serving path tries them from most specific to most general and stops at the first eligible current product.
That hierarchy captures the offline-online split cleanly. The offline model decides which need from the consumer’s history is worth carrying forward; the online system decides which product can satisfy that need right now.
Before a bundle is published, deterministic validation sits between model generation and the serving contract. It distinguishes an explicit abstention from malformed output, removes unusable identifiers, normalizes quantities, and filters results that do not meet the requirements for a useful bundle. The model makes the semantic decision; deterministic code decides whether the result is safe to serve.
Generating and serving at scale
Generating these bundles across a large consumer cohort requires a batch system that can tolerate partial failures. We use a sharded Metaflow infrastructure so batches can run independently and failed shards can be retried without discarding successful work.
After validation, the resulting bundles are published to a low-latency key-value store. The serving path can then fetch a consumer’s bundle for a given store without repeating the historical reasoning, while missing or unusable bundles degrade to an empty response rather than failing the Grocery turn.
Because the bundle is generated independently of Ask, the same signal can also support other personalized surfaces without repeating the inference.
None of this is visible to the consumer. What they see is a shopping list — and depending on how they entered the flow, they may see it in about two seconds.
User Experience: One list, several ways in
A consumer does not always arrive at reorder through the same path. They may type “reorder my usuals” directly into Ask, or they may enter through a contextual surface where DoorDash already knows something about what they are trying to do.
That distinction matters because context supplied by the experience can remove work from the conversation. As described in Building Ask DoorDash (Part 5): A Grounded Interface for Shopping Agents, the client attaches scope to a request, carrying information such as the topic and store the consumer entered from. The system can use that context instead of asking the consumer to restate information the product already knows.
For a free-form request like “reorder my usuals” the Grocery agent still interprets the request through the Reorder skill. But when a consumer enters through a dedicated “Reorder my usuals” action, the product has already established the intent. As shown below, we attach a preset scope that is invisible to the consumer and route the request through a deterministic path.

Combined with the precomputed bundle from the Platform section, that reduces the live work to retrieving the bundle and resolving its items against the current catalog. The path no longer needs the LLM to interpret the request or reconstruct the consumer’s usuals from history, bringing latency from roughly 16 seconds to roughly 2 seconds.
Whether the consumer arrives through free-form text or a contextual entry point, the result is the same grounded shopping-list artifact. From there, they can change quantities, remove something they already have, swap a product, or add the finished list to their cart. Those structured edits happen directly against the artifact rather than requiring another LLM round trip.
The agent remains available when the request requires more judgment. A consumer can ask for “my usuals, but cheaper” for example, and the agent can start from the same recurring needs while reasoning over current prices and alternatives.
Different entry points can trigger very different amounts of work underneath while still converging on one continuous reorder experience.
Conclusion
“Reorder my usuals” looks like a simple request, but making it useful requires putting different decisions in the right place. Consumer context helps interpret purchase history; recurring needs can be computed ahead of time; inventory, prices, and catalog resolution stay live; and contextual entry points can remove work the product has already done.
The broader lesson is that making an agent better does not always mean asking the model to do more. It means deciding what should be reasoned about at request time, what can be known ahead of time, and what should be deterministic.
When those boundaries are right, the complexity disappears for the consumer. They ask DoorDash to reorder their usuals and get a list that feels personal, current, and ready to shop.
Stay Informed with Weekly Updates
Subscribe to our Engineering blog to get regular updates on all the coolest projects our team is working on
