Grocery Journey
Recurring shopping creates opportunities to reduce effort, but only specific conditions justify autonomous action. Preparation, approval, execution, and confirmed purchase are different states.
Repetitive household grocery shopping is not just my problem. Using emerging agentic commerce technologies, I researched, designed, and built a working agent end to end — including delegated permission to spend and payment execution. The result was a verified autonomous purchase in a controlled environment.
See what actually worksProduct strategy · Research · UX · Systems design · PrototypingImplementation · Testing & verification · Platform: Android · Web · API
R1–R8Deterministic purchase authorization and exception rules
3 decisionsACT · ASK · STOP, based on delegated household authority
$3.69 TESTOne controlled autonomous milk purchase · September 21, 2026
A memory-first grocery agent that remembers household preferences and confirmed purchases, evaluates purchasing rules, and determines when it can act independently or needs human approval.
Recurring shopping creates opportunities to reduce effort, but only specific conditions justify autonomous action. Preparation, approval, execution, and confirmed purchase are different states.
Remembering what is running low offers value even to people who do not want automatic purchasing.
Agentic commerce is moving fast. Discovery and checkout are becoming agentic, while payments still depend on trust, explicit authorization, and clear accountability.
A cart is not a completed purchase. “How autonomous should the agent be?” “What authority should the shopper be able to delegate—and how should the agent behave within it?”
I mapped the repeated tasks and judgment calls involved in keeping a household stocked—from noticing a need to completing a purchase.
Is this actually needed now?
Is this still the right product and quantity?
Do today’s conditions still qualify?
Is the purchase within household boundaries, and did it actually go through?
Directional findings from 13 respondents; these results describe this small sample, not consumers generally.
| Survey finding | Responses (n=13) | Product implication |
|---|---|---|
| Remembering what is running low was effortful | 9 of 1369.2% | Make household memory and replenishment suggestions useful independently of buying. |
| Shopped primarily in-store | 11 of 1384.6% | Support planning even when the eventual purchase does not happen inside the agent. |
| Would consider rule-bound auto-buying of household supplies | 6 of 1346.2% | Make routine purchasing explicitly authorized rather than an assumed default. |
| Did not want AI purchasing for them | 5 of 1338.5% | Allow people to use memory and planning without granting purchase authority. |
| Chose fresh produce, meat/seafood, or expensive specialty items for auto-buying | 0 of 130% | Use product/category boundaries and require explicit decisions for sensitive purchases. |
I compared public capabilities across shopping assistants, retailer agents and payment protocols. The strongest market signal is that discovery and checkout are becoming agentic. This comparison focuses on consumer-facing products; UCP, ACP, and AP2 are commerce and authorization mechanisms, rather than consumer-agent competitors.
| Existing approach | What I investigated | Buyer’s Agent question |
|---|---|---|
| Shopping assistance, reordering and selected automated purchase capabilities within Amazon’s shopping environment. | How could a household define standing purchasing rules that are not limited to one retailer account? | |
| Conversational product assistance and shopping journeys within Walmart’s ecosystem. | Could household preferences and approval boundaries persist across different merchant environments? | |
| Natural-language grocery discovery and cart preparation across the platform’s supported retailer network. | What would household-level permission look like before a recurring purchase—not only at cart or checkout? | |
| Personal context and assisted checkout within a platform and connector environment. | Could household-defined purchasing authority remain portable when shopping and checkout capabilities come from platform connectors? | |
| Interoperable commerce primitives plus verifiable intent and payment mandates. | How could household policy decisions generate or consume standardized authorization evidence before execution? | |
| Agentic checkout, merchant order lifecycle and shared payment primitives. | What should determine whether the agent is allowed to act before transaction execution begins? | |
| Visa / MastercardPayment-network controls | Agent identity, verifiable intent, scoped credentials and payment-network trust controls. | How can a human household rule become an explainable decision before a payment rail is invoked? |
A dated research comparison, not a claim that retailer agents do not benefit consumers or that no other buyer-directed agents exist. Linked product descriptions provide the reference points.
Close research details ↑The goal was not maximum autonomy. It was matching agent autonomy to the authority the shopper was willing to delegate.
| Autonomy level | Delegation level | What the agent does | What the human does | Current retail / commerce state |
|---|---|---|---|---|
| Assist | Low · Task execution | Searches, compares, summarizes, and answers shopping questions. | Defines the need, evaluates options, and purchases. | Established |
| Recommend | Low–moderate · Task execution | Suggests products, alternatives, or possible replenishment needs. | Reviews the recommendation and decides whether to act. | Established |
| Prepare | Moderate · Task execution | Selects products, quantities, or merchants and prepares a cart or checkout. | Reviews the proposal and completes or approves the purchase. | Established / expanding |
| Decide within boundaries | Higher · Delegated decisions | Chooses among acceptable options using preferences, budget, and other constraints. | Defines the boundaries and reviews or approves the resulting transaction. | Emerging |
| Act within boundaries | High · Autonomous purchasing | Initiates and completes an eligible purchase without asking every time. | Grants authority in advance, sets limits, monitors activity, and can revoke permission. | Early / bounded |
| Open autonomy | Very high · Broad delegation | Determines need, chooses products, and purchases with little predefined constraint. | Has minimal involvement. | Not broadly established |
Reconstructed example—not a verbatim prompt from my original chat: “Compare grocery assistants by what they remember, whose catalog they use, whether they suggest, prepare a cart, or complete a purchase, and when approval is required. Separate verified capabilities from assumptions and cite the original product sources.”
These are research methods, not a claim that a named reusable AI skill was installed. The original prompt wording has not been verified.
Tools across discovery and feasibility research: ChatGPT · Perplexity · Claude. AI supported research preparation and synthesis; the survey responses and product decisions remained mine.
when should a household agent act on someone's behalf? Read my research above
Buy again from purchase history or saved lists.
Subscriptions or repeat purchases.
Build a cart with past items or AI suggestions.
AI helps search, choose and move to checkout.
Auto-Buy &
Reorder,
Household
MCP
Muse
Sarah, household buyerWhat a household usually buys.
A signal, not certainty, that an item may be running low.
Permission, limits and conditions set by the household.
Remembers household preferencesUses purchase evidence and corrected household information.
Identifies possible replenishment needsUses available household signals without assuming current stock.
Evaluates rules and conditionsChecks the applicable grant, product, quantity, price, merchant and cadence.
Executes eligible routine purchasesOnly when need, permission and merchant support align; otherwise asks or stops.
I framed the experience around the household buyer who owns the decision and Buyer’s Agent, which can interpret context and act only inside delegated authority.
Reduce repetitive shopping work while staying in control of what gets bought.
Repeated decisions, uncertain replenishment needs, and concern about unwanted autonomous purchases.
Interpret household context, propose actions, check authority, and act only within explicit boundaries.
It can infer a need, but it cannot give itself permission to buy.
Routine purchasing with clear boundaries and less interruption.
A household staple may be running low.

Buyer’s Agent suggests a specific product.

1 × 1L carton
$4.49Buyer’s Agent checks household rules and decides.
Buyer’s Agent completes the purchase and confirms the result.
Key interface moments across the journey.
Proactive insight about household needs.
See the reasoning behind suggestions.
Review and approve when needed.
See the verified outcome and delivery progress.
Buyer’s Agent recognizes that a household staple is due for replenishment.

The agent already knows the exact item, quantity, and preferred merchant.

1 × 1 carton (52 oz)
$3.69
Buyer’s Agent checks the standing household rule and confirms it can act.
Buyer’s Agent places the order and records the result.
Key interface moments across the journey.

Household sets and saves their preferred item, quantity, and merchant.
Household grants ongoing authorization for eligible routine purchases.

Agent completes the purchase and records the transaction details.
Household can view status and delivery progress anytime.
Controlled evidence: One owner-operated Market WooCommerce sandbox purchase of Fairlife milk (quantity 1, $3.69) reached ORDERED with a successful Stripe TEST payment on September 21, 2026. The need signal was synthetic; order #2093 was processing. The original concept image’s timing, shipping estimate, and saved-payment implication are illustrations—not independently verified household or fulfillment evidence.
Open-ended exploration of the household problem, autonomy, trust, and possible experiences.
Generate HTML, critique flows, compare alternatives, and refine interaction and agent rules.
Product rules, system states, design constraints, visual conventions, and implementation instructions.
Scope, flows, constraints, and acceptance criteria.
Interaction semantics, authority language, and state meaning.
ACT / ASK / STOP, lifecycle, and explicit system boundaries.
Durable decisions, version history, and implementation guidance.
Product intent becomes testable behavior.
Make interactions, journeys, states, handoffs, and system boundaries visible enough to judge.
Preserve interaction and visual intent.
Define intent, authority, and state semantics.
Preserve durable product and architecture decisions.
Make expected behavior testable.
Give coding agents a bounded operating context.
Confirmed purchases, familiar items, quantities, cadence, merchant preference, and explicit instructions.
Interpret a request or infer that a familiar staple may be running low.
Product identity, merchant, price, availability, quantity, and checkout capability.
Check product, merchant, price, familiarity, substitution, cadence, category, and grant.
Recheck final conditions, claim one attempt, and invoke the supported checkout/payment rail.
Confirm whether an order exists before the interface claims success.
No grant → ASK rather than invent authority.
Hard restriction → STOP; unresolved condition → ASK.
A different merchant can require approval.
Current price or final total above the grant → ASK.
A new product cannot inherit authority from a familiar staple.
OOS replacement must respect explicit substitution rules.
All conditions inside the household grant → ACT.
If the engine cannot resolve safely, fall back to ASK.
Present household context and agent state to the buyer.
Maintain the evidence and household context the system can use.
Decide whether the proposed action is permitted.
Translate an authorized action into the capability each merchant exposes.
Determine what actually happened after an authorized attempt.
Live merchant offer, deterministic ACT decision, guarded execution, TEST payment, merchant order evidence, and reconciliation to ORDERED.
Discovery, catalog lookup, cart creation, and cart reread were verified. Autonomous checkout and payment permission were not granted.
Potential future rails for standardized authorization evidence, interoperable checkout, and production payment capability. Not part of the current verified execution proof.
| Layer | Technology / protocol | Case-study status |
|---|---|---|
| Household policy | Buyer’s Agent R1–R8 | Built |
| Data + memory | Supabase | Built |
| Merchant execution | WooCommerce APIs | Controlled path verified |
| Payment | Stripe TEST | TEST payment verified |
| Agent commerce | Shopify UCP / MCP | Cart-ready verified |
| Authorization standard | AP2 | Research / future direction |
| Agentic checkout | ACP | Research / future direction |
| Household policy | Buyer’s Agent | Built / proprietary core |
| Authorization evidence | Google AP2 | Explored · not implemented |
| Merchant interoperability | Shopify / Google UCP | Cart-ready path verified |
| Agentic checkout | OpenAI + Stripe ACP | Explored · not implemented |
| Payment credentials | Stripe SPT / Visa Intelligent Commerce | Explored · not implemented |
| Agent identity / merchant trust | Visa / Mastercard protocols | Explored · not implemented |
State the desired behavior and exact product boundary.
Design docs, authority policy, contracts, and relevant ADRs.
Use coding agents to accelerate implementation against existing architecture.
Does the implementation preserve authority, state, and interaction semantics?
Tests, merchant behavior, safeguards, and actual outcome.
Commit code and update durable context when product truth changes.
Foundational design logic.
Standard implementation components where they improve consistency and maintainability.
Product meaning that the foundation library does not provide.
Household scenario: Sarah’s household may be running low on milk. Buyer’s Agent checks what she usually buys and the household’s current purchasing authority to decide whether it can ACT, needs to ASK, or must STOP.
The agent needs human judgment.
This action falls outside the current grant. Nothing is ordered yet.

The agent has authority and is acting.
The action is authorized. The UI does not yet claim an order exists.
The agent does not yet know the outcome.
The outcome is unknown. Buyer’s Agent does not retry while the existing attempt is unresolved.
Verified evidence allows the product to claim success.
The store confirmed the order and the sandbox TEST payment succeeded.

The agent needs human judgment.

The agent has authority and is acting.

The agent does not yet know the outcome.

Verified evidence allows the product to claim success.

Household scenario: Sarah’s usual milk is inside the household’s standing grant. Buyer’s Agent checks the current offer and authority, resolves to ACT, and begins the purchase without asking again. ACT gives the agent permission to execute. It does not yet prove that an order exists.
The agent is allowed to act within those household instructions.
The household has granted standing authority for routine items.
The grant defines product, price, store, cadence, and substitution boundaries.
Verified order evidence shows that execution completed successfully.
State alone was not enough. V4 proved the working execution model. The system protects against duplicate attempts, final-total changes, missing responses, and authority changes while the transaction is in progress.
Does the household actually need this now? Purchase history can suggest a possible need, not prove one.

Can the usual product still be bought under the expected conditions? Availability, substitutions, price, and merchant conditions can change.

Do the current conditions still fall inside the household’s permission? Knowing what to buy is not permission to buy it.

Did the purchase actually happen? A missing or delayed response is not proof of failure; reconcile the existing attempt before retrying.

Uncertainty can enter before a purchase, at the live offer, at the authority boundary, or after execution begins.
A replacement can be found without silently inheriting the original item’s authority.

If the live price moves outside the household grant, execution stops or returns the decision to the person.

A pause or revocation changes what the agent may do next, even while checkout is underway.

A missing response does not prove failure. Preserve the attempt and reconcile before any retry.

User scenario: Sarah’s household may be running low on milk. What needs her attention right now? Shopping, agent permissions, household context?
Shopping List, merchant comparison, checkout routes, and payment led the experience.

Permissions and buying instructions became the landing experience.

Memory, routines, and product status became primary.

Start with what needs attention now; keep memory and authority one layer underneath.

User scenario: Sarah wants Buyer’s Agent to keep routine groceries stocked, but she should not have to understand grants, policy rules, or ACT / ASK / STOP. When the agent reaches a decision, how should the experience tell her whether she needs to act, the agent is working, or the task is complete?
Capability was presented as the experience.

Level 2, Level 3, purchase mandate, and situational authority exposed the architecture.

Policy outcomes were surfaced directly as product language.

Needs your decision · Buying · Checking order status · Bought for you.

User scenario: Buyer’s Agent has permission to buy Sarah’s milk and begins the purchase, but the merchant response may be delayed, incomplete, or uncertain. When is the interface allowed to tell Sarah that the milk has actually been bought?
The product implied the task was complete as soon as the agent decided to act.

Intent and purchase outcome were compressed into one user-facing statement.

The state model improved, but authorization still sat too close to transaction completion.

Execution, unresolved outcome, and verified completion each get a distinct human-facing state.



I tested policy decisions, execution safeguards and controlled commerce independently. The milestone matters because it connects a permitted action with an actual merchant order, successful Stripe TEST payment and a reconciled system state—not simply a convincing screen.
A correct product suggestion is not enough. The proposal, authority decision, tool execution, and final UI all need to agree with the underlying evidence.
Check product identity, quantity, household context, ambiguity, and whether the system asked for clarification when evidence was weak.
Compare the scenario against the deterministic rule contract, applicable grant, price ceiling, merchant, substitution, and cadence.
Verify supported merchant path, one-use attempt, final-total recheck, quantity, idempotency, and reconciliation behavior.
Ensure “authorized,” “buying,” “checking,” “failed,” and “bought” match the actual execution and order evidence.
The existing automated suites primarily test deterministic policy and execution behavior. They are not a complete measure of LLM quality. A versioned scenario set for ambiguous household requests and measured interpretation performance is a next-stage evaluation task.
A correct recommendation is not enough. I evaluate the model interpretation, deterministic authorization, tool behavior, and the truthfulness of the final interface as separate layers.
Did the proposal match the household’s product, quantity, and known context? Did ambiguity trigger clarification?
Did the scenario produce the expected ACT / ASK / BLOCK result and the correct user-facing authorization state?
Did the agent stay on a supported merchant path, respect the one-use gate, recheck the total, and avoid duplicate execution?
Did the explanation and status match the actual rule, execution state, and order evidence?
I separate what the product has actually executed from what remains access-dependent. The goal is truthful evidence, not a demo that implies unsupported merchant or payment capability.
One authorized Fairlife milk purchase reached RECONCILING → ORDERED with successful Stripe TEST payment evidence.


Catalog lookup, cart creation, and cart retrieval were verified. Checkout, payment, and a Shopify order were not executed through UCP.


Buyer’s Agent made an ACT / R7 decision for one Fairlife lactose-free 2% milk, executed one authorized purchase against the controlled Market WooCommerce store, and independently verified the sandbox payment and order.
The household signal used a synthetic low-stock fixture, not a real family’s pantry sensor or an immutable historical household snapshot. This was the controlled owner-operated store and Stripe TEST mode: no goods were packed, shipped or delivered. Read the sanitized transaction evidence .
Project notes record V4 runs of 65 pipeline checks, 39 execution-store checks and 28 ceiling checks. These are separate suite-level counts and must not be summed as distinct tests or presented as current production pass-rate metrics without fresh run logs.
In a separate owner-controlled Shopify development store, I discovered the merchant UCP endpoint and used catalog lookup, create_cart and get_cart for one Kitchen Sponge. The same cart was reread with quantity one, estimated total $2.99 and a preserved continuation URL.
Not executed: UCP checkout creation, checkout completion, payment or Shopify order placement. The development-store test gateway and advertised tools do not grant delegated checkout authority.
For each layer, I distinguish what has been tested from the measurement that would make it credible in production.
| Layer | Current evidence | Still to measure |
|---|---|---|
| Need interpretation | Representative requests, synthetic household fixtures and prototype review. | Versioned ambiguous-request dataset; product/quantity matching, clarification and unsupported-claim rates. |
| Deterministic authority | Documented R1–R8, ceiling and execution-store test coverage. | Reproducible expected-versus-actual scenarios and regression results after every rule change. |
| Tool execution | One verified WooCommerce TEST order; single Shopify UCP cart-create-and-read milestone. | Cross-merchant reliability, unauthorized-tool-attempt rate and approved payment interoperability. |
| Experience truth | Reviewed ASK-card anatomy, distinct order states and local/device prototype checks. | User comprehension of AUTHORIZED versus ORDERED, approval burden, correction success and perceived control. |
Established: a bounded policy can authorize and execute one controlled TEST-store purchase with verified order/payment evidence, and a separate Shopify UCP path can prepare a cart. Not established: autonomous purchasing for arbitrary retailers, real-money delegated payment, household adoption, long-term accuracy or reliable fulfillment. Those require merchant access, payment permissions, security review and real-user validation.
I began by asking how much grocery work an AI assistant could take on. Building the system changed the question: how can a household understand the agent’s memory, specify its purchasing authority, and know whether a transaction actually occurred?
A useful assistant can remember familiar products, surface likely needs and prepare a decision without permission to spend. That gives households a meaningful, lower-stakes way to understand the system first.
AUTHORIZING a purchase, ATTEMPTING it and VERIFYING an order each deserve a different state, explanation and recovery path. A fluent answer cannot substitute for a merchant receipt.
Documenting the decision contract, UI semantics and execution lifecycle let me challenge AI-generated work against explicit expectations. The value of a design system here is behavioral consistency as much as visual consistency.
The controlled WooCommerce + Stripe TEST purchase reached ORDERED; Shopify UCP catalog/cart operations were independently verified without payment. The V4 Android closed-testing build and V5 interactive prototype demonstrate different stages of experience readiness, not a released real-money purchasing service.
Live household purchasing needs authorized retailer checkout access, delegated-payment controls, secure credential handling, customer consent, order reconciliation, real-world failure testing and release review. A working demonstration on a store I control does not grant those capabilities across other retailers.
Observe how people correct memory, set specific grants, understand ASK versus ORDERED and revise their delegation over time.
Measure replenishment precision, missed/early suggestions, time saved, approval burden and trust after repeated shopping cycles.
Validate allowed merchant/payment integrations, final-total rechecks, concurrent attempts, timeout reconciliation, refunds and operational monitoring under controlled conditions.
Build a versioned scenario set for ambiguous requests and model suggestions; measure interpretation errors separately from policy and execution failures.
My guiding question remains: How much purchasing authority should an AI agent actually have—and how can a person see, change and revoke that authority?
The current project proves controlled system behavior, not real-household impact. The next measure of success is whether Buyer’s Agent makes routine shopping meaningfully easier while people can still understand, correct and revoke what it is allowed to do.
Measure time spent on repeat purchases, number of manual steps, suggestion acceptance versus correction, and whether the agent reduces unnecessary shopping decisions across repeated grocery cycles.
Measure comprehension of AUTHORIZED, ASK and ORDERED states; success editing or revoking grants; approval burden; correction success; and perceived control after the agent acts.
Measure policy mismatches, unauthorized-action attempts, checkout success on supported merchants, reconciliation accuracy, duplicate-attempt prevention, and recovery from ambiguous or incomplete transaction states.
Today I can point to one controlled WooCommerce + Stripe TEST purchase and one Shopify UCP cart-create-and-read milestone. Those prove specific system capabilities; they do not establish time saved, adoption, trust, replenishment accuracy or long-term household value.
Run a longitudinal household pilot with a pre-agent baseline, then track repeated shopping cycles: what the agent inferred, what people corrected, when it asked, when it acted, whether the order was verified, and how much work the household still had to do.
Success is not “the agent bought something.” Success is fewer routine decisions and less shopping effort without hidden authority, unwanted purchases or uncertainty about what actually happened.
Merchant offers, model output and tool responses are treated as inputs to evaluate — not as permission. Consequential actions stay behind household policy and supported execution paths.
Versioned ambiguous-request scenarios and regression grading.
End-to-end visibility across model interpretation, tools, policy and outcomes.
Prompt-injection, malformed tool data and external-content attacks.
Turn real failures into new eval cases before the next release.
Design semantics, engineering instructions, architecture decisions and execution safeguards are versioned alongside the implementation.