Agentic Commerce - Autonomous Buying

Buyer’s Agent app iconHow much purchasing authority should an AI agent actually have?

Repetitive household grocery shopping is not just my problem. Using emerging agentic commerce technologies, I researched, designed, and built a working agent end to end — including delegated permission to spend and payment execution. The result was a verified autonomous purchase in a controlled environment.

See what actually works
Role
Independent Product Designer & Builder ·
Status
Working sandbox demo (as of Oct 2026).
Successful autonomous purchase verified with WooCommerce & Stripe test payment; Ready to check out experience with Shopify UCP.
What I found
Production autonomous purchasing requires authorized retailer integration, customer-approved payment access, and verified order execution.

Meet Buyer’s Agent.

A memory-first grocery agent that remembers household preferences and confirmed purchases, evaluates purchasing rules, and determines when it can act independently or needs human approval.

02 · Discovery

I researched:household grocery habits, user needs, emerging commerce agents, and the boundaries of autonomous purchasing.

CURRENT STATE

Grocery Journey

What I learned

Recurring shopping creates opportunities to reduce effort, but only specific conditions justify autonomous action. Preparation, approval, execution, and confirmed purchase are different states.

USER RESEARCH

Household Survey

What I learned

Remembering what is running low offers value even to people who do not want automatic purchasing.

EMERGING LANDSCAPE

Commerce & Competitors

What I learned

Agentic commerce is moving fast. Discovery and checkout are becoming agentic, while payments still depend on trust, explicit authorization, and clear accountability.

AGENTIC SHOPPING

Autonomy & Delegation

What I learned

A cart is not a completed purchase. “How autonomous should the agent be?” “What authority should the shopper be able to delegate—and how should the agent behave within it?”

Current state · Household shopping journey

Routine grocery shopping mixes :in-store trips, online reorders, subscriptions, and assisted checkout.

I mapped the repeated tasks and judgment calls involved in keeping a household stocked—from noticing a need to completing a purchase.

Shopper actionsHousehold tasks
Notice a possible needRemember what is running low and when it may be needed.
Choose what to buyPick the usual brand, size, quantity, and acceptable alternatives.
Evaluate today’s conditionsCheck current price, availability, timing, and substitutes.
Approve and complete purchaseApprove spend, place the order, and recognize whether it succeeded.
1

Notice

Digital support used today
  • Shopping lists
  • Reorder history
  • Reminders
Judgment still required

Is this actually needed now?

2

Choose

Digital support used today
  • Saved items
  • Reorder tools
  • Scheduled buys / subscriptions
Judgment still required

Is this still the right product and quantity?

3

Evaluate

Digital support used today
  • Price tracking
  • Retailer suggestions
  • Cart preparation
Judgment still required

Do today’s conditions still qualify?

4

Purchase

Digital support used today
  • Checkout
  • Scheduled cart add
  • Agent-assisted checkout
Judgment still required

Is the purchase within household boundaries, and did it actually go through?

Close research details ↑
User research · Exploratory household survey · July 2026

Real buyers: families of 3 and above, busy households with routine grocery shopping.

Directional findings from 13 respondents; these results describe this small sample, not consumers generally.

Survey findingResponses (n=13)Product implication
Remembering what is running low was effortful
9 of 1369.2%
Make household memory and replenishment suggestions useful independently of buying.
Shopped primarily in-store
11 of 1384.6%
Support planning even when the eventual purchase does not happen inside the agent.
Would consider rule-bound auto-buying of household supplies
6 of 1346.2%
Make routine purchasing explicitly authorized rather than an assumed default.
Did not want AI purchasing for them
5 of 1338.5%
Allow people to use memory and planning without granting purchase authority.
Chose fresh produce, meat/seafood, or expensive specialty items for auto-buying
0 of 130%
Use product/category boundaries and require explicit decisions for sensitive purchases.
Close research details ↑
Commerce & competitor landscape · 2026 research snapshot

Shopping is changing fast. Who helps the buyer and what defines the purchasing boundaries?

I compared public capabilities across shopping assistants, retailer agents and payment protocols. The strongest market signal is that discovery and checkout are becoming agentic. This comparison focuses on consumer-facing products; UCP, ACP, and AP2 are commerce and authorization mechanisms, rather than consumer-agent competitors.

Existing approachWhat I investigatedBuyer’s Agent question
Amazon Alexa for Shopping ↗Retailer-owned shoppingShopping assistance, reordering and selected automated purchase capabilities within Amazon’s shopping environment.How could a household define standing purchasing rules that are not limited to one retailer account?
Walmart Sparky + ChatGPT / Gemini ↗Retailer-owned shoppingConversational product assistance and shopping journeys within Walmart’s ecosystem.Could household preferences and approval boundaries persist across different merchant environments?
Instacart + ChatGPT / Gemini ↗Grocery-platform shoppingNatural-language grocery discovery and cart preparation across the platform’s supported retailer network.What would household-level permission look like before a recurring purchase—not only at cart or checkout?
Meta · MusePersonal agent / connector contextPersonal context and assisted checkout within a platform and connector environment.Could household-defined purchasing authority remain portable when shopping and checkout capabilities come from platform connectors?
Google UCP + AP2Commerce & authorization infrastructureInteroperable commerce primitives plus verifiable intent and payment mandates.How could household policy decisions generate or consume standardized authorization evidence before execution?
OpenAI + Stripe ACPAgentic checkout infrastructureAgentic checkout, merchant order lifecycle and shared payment primitives.What should determine whether the agent is allowed to act before transaction execution begins?
Visa / MastercardPayment-network controlsAgent identity, verifiable intent, scoped credentials and payment-network trust controls.How can a human household rule become an explainable decision before a payment rail is invoked?

A dated research comparison, not a claim that retailer agents do not benefit consumers or that no other buyer-directed agents exist. Linked product descriptions provide the reference points.

Close research details ↑
Agentic Shopping · Autonomy & Delegation

How much of the shopping journey can an agent take on?

The goal was not maximum autonomy. It was matching agent autonomy to the authority the shopper was willing to delegate.

AutonomyDescribes how independently the agent can act.
DelegationDescribes how much authority the shopper has intentionally given the agent.
Autonomy level Delegation level What the agent does What the human does Current retail / commerce state
AssistLow · Task executionSearches, compares, summarizes, and answers shopping questions.Defines the need, evaluates options, and purchases.Established
RecommendLow–moderate · Task executionSuggests products, alternatives, or possible replenishment needs.Reviews the recommendation and decides whether to act.Established
PrepareModerate · Task executionSelects products, quantities, or merchants and prepares a cart or checkout.Reviews the proposal and completes or approves the purchase.Established / expanding
Decide within boundariesHigher · Delegated decisionsChooses among acceptable options using preferences, budget, and other constraints.Defines the boundaries and reviews or approves the resulting transaction.Emerging
Act within boundariesHigh · Autonomous purchasingInitiates and completes an eligible purchase without asking every time.Grants authority in advance, sets limits, monitors activity, and can revoke permission.Early / bounded
Open autonomyVery high · Broad delegationDetermines need, chooses products, and purchases with little predefined constraint.Has minimal involvement.Not broadly established
Close research details ↑
Read about: AI-assisted discovery,how AI helped me
Survey design & synthesisExplore questions, group response patterns, and distinguish what participants said from assumptions.
Commerce researchBuild a consistent comparison of buyer benefit, merchant context, shopping capability, and purchase authority.
Autonomy critiqueProbe scenarios where an agent should remember, recommend, prepare, ask, or act.
See a research prompt example

Reconstructed example—not a verbatim prompt from my original chat: “Compare grocery assistants by what they remember, whose catalog they use, whether they suggest, prepare a cart, or complete a purchase, and when approval is required. Separate verified capabilities from assumptions and cite the original product sources.”

These are research methods, not a claim that a named reusable AI skill was installed. The original prompt wording has not been verified.

Tools across discovery and feasibility research: ChatGPT · Perplexity · Claude. AI supported research preparation and synthesis; the survey responses and product decisions remained mine.

Discovery synthesis

Shopping automation is evolving,but the household question is more than just making a purchase.

when should a household agent act on someone's behalf? Read my research above

01
What already exists

Shopping automation spans the journey.

Reorder tools

Buy again from purchase history or saved lists.

Scheduled reordering

Subscriptions or repeat purchases.

Cart preparation

Build a cart with past items or AI suggestions.

Agent-assisted checkout

AI helps search, choose and move to checkout.

Examples (illustrative)
AmazonAuto-Buy &
Scheduled Actions
WalmartReorder,
Subscriptions & Lists
InstacartHousehold
shopping & delivery
ShopifyMCP
AI-assisted commerce
MetaMuse
Personal AI agent
(browser checkout)
02
The household buyer question

Automating steps does not settle the decision.

Do we need this item now?
Sarah, household buyer, considering two questionsSarah, household buyer
Does the agent have permission to buy it today?
Need and authority are two separate decisions.
03
The distinction

Three different things, often confused.

Historical preference

What a household usually buys.

Past shopping
data and habits.
Possible current need

A signal, not certainty, that an item may be running low.

Depleted pantry
items, consumption
patterns, context
and timing.
Purchasing authority

Permission, limits and conditions set by the household.

Price limits, sensitive
items, cadence rules
and other policies.
04
The opportunity

A buyer-directed household agent.

Illustration of Buyer’s Agent holding a grocery basket

Remembers household preferencesUses purchase evidence and corrected household information.

Identifies possible replenishment needsUses available household signals without assuming current stock.

Evaluates rules and conditionsChecks the applicable grant, product, quantity, price, merchant and cadence.

Executes eligible routine purchasesOnly when need, permission and merchant support align; otherwise asks or stops.

Who I designed for

I framed the experience around the household buyer who owns the decision and Buyer’s Agent, which can interpret context and act only inside delegated authority.

Primary user · Sarah · Household buyer Sarah, household buyer
Goal

Reduce repetitive shopping work while staying in control of what gets bought.

Friction

Repeated decisions, uncertain replenishment needs, and concern about unwanted autonomous purchases.

System actor · Buyer’s Agent Buyer’s Agent
Responsibility

Interpret household context, propose actions, check authority, and act only within explicit boundaries.

Constraint

It can infer a need, but it cannot give itself permission to buy.

The intended household buyer experience

Routine purchasing with clear boundaries and less interruption.

Set boundariesDefine the usual item, quantity, merchant, price limits, substitutions, and other purchasing conditions.
Let routine work happenThe agent can act when the need qualifies and the purchase stays within those boundaries.
Stay informedSee what the agent did, why it acted, and the verified transaction outcome.
Step in when neededMake a decision when the purchase falls outside the household’s existing authority.
03 · Core Journey

Core User Journey

Shopping experience · Household Buyer & Buyer’s Agent
A · Human Review & Approval Path

Buyer Approved Purchase

1

Need identified

A household staple may be running low.

Warm kitchen scene
TodayHousehold overview
Oat milk may be running lowPossible need from purchase history · confirm first
2

Product proposed

Buyer’s Agent suggests a specific product.

Oatly Oat Milk Original carton
Oatly Oat Milk Original

1 × 1L carton

$4.49
Whole Foods Market logo
Whole Foods MarketIn stock · Good price
✓Looks good
See alternatives
3

Policy evaluated

Buyer’s Agent checks household rules and decides.

Checking household rules…
ACTMatches an approved category (“Everyday Staples”)
ASKFirst time buying this brand from this merchant
BLOCKNot blocked by any rules (price, merchant, or category)
Decision: ASK
This is a new brand at a new merchant. I’ll ask for your approval.
4

Purchase attempt → verified outcome

Buyer’s Agent completes the purchase and confirms the result.

Attempting purchase…Checking out with Whole Foods Market
Purchase successfulOrder #WF3287619
Verified outcome1 item purchased · $4.49 · Arriving Tue, Apr 23
Shopper actions
Notice and confirm need“Yep, we’re running low.”
Review when asked“This looks good.”
Provide approval (if needed)“Yes, go ahead.”
View outcome“Great, it’s ordered!”
Agent actions
Detect possible needMonitors household usage and predicts a potential need.
Propose productSelects a suitable item, quantity, and merchant.
Evaluate rulesChecks household policies and decides (ACT / ASK / BLOCK).
Prepare purchase → attempt checkoutCompletes purchase and reconciles the outcome.

Experience touchpoints

Key interface moments across the journey.

TodayMon, Apr 22
Oat milk may run out soon
Today

Proactive insight about household needs.

Why this suggestion?×
◉ Based on your usage (~1 carton per week)◉ You usually buy oat milk
View why

See the reasoning behind suggestions.

Approve purchase?Oatly Oat Milk Original · 1 × 1L · Whole Foods
Not nowApprove
Approval

Review and approve when needed.

✓
Order placedArriving Tue, Apr 23
Ordered
Packed
On the way
Delivered
Order status

See the verified outcome and delivery progress.

Tested commerce path

Need identified → Product found → Added to cart → Merchant checkout unavailable → Human handoff

B · Standing Authorization Path · Verified TEST milestone

Routine autonomous purchase by Buyer Agent

1

Routine need detected

Buyer’s Agent recognizes that a household staple is due for replenishment.

Warm kitchen scene
TodayHousehold overview
Fairlife milk running lowBased on your routine purchase pattern and household preference.
2

Preference already known

The agent already knows the exact item, quantity, and preferred merchant.

Fairlife Lactose-Free 2% Milk carton
Fairlife Lactose-Free 2% Milk

1 × 1 carton (52 oz)

$3.69
Market logo
Market (Demo Store)Preferred store
Saved preference
Standing permission
No comparison neededUsing your saved preference.
3

Policy qualifies purchase

Buyer’s Agent checks the standing household rule and confirms it can act.

Checking standing authorization…
ROUTINEHousehold staple
ITEMMatches saved preference (Fairlife 2% Milk)
PRICEWithin budget ($3.69 ≤ $5.00)
MERCHANTApproved merchant (Market)
CADENCELast purchased 12 days ago (next due ~14 days)
Decision: ACT
This purchase matches the household’s saved preference and standing authorization. No approval needed.
4

Purchase completed → verified outcome

Buyer’s Agent places the order and records the result.

Attempting purchase…Checking out with Market
Purchase successfulOrder #2093
Verified outcome1 carton purchased · $3.69 · WooCommerce + Stripe TEST · Arriving Thu, Apr 25
Shopper actions
Set preference earlier“Buy my usual milk when it runs low.”
Grant standing permission“Use my saved preferences.”
No action needed during purchase“I don’t need to do anything.”
Review outcome later“Great, it’s handled.”
Agent actions
Detect routine needMonitors household usage and predicts replenishment.
Use saved preferenceRetrieves the exact item, quantity, and merchant.
Check policyValidates against standing rules and authorization.
Prepare purchase → attempt checkoutPlaces the order automatically using saved credentials.
Reconcile and record outcomeConfirms purchase and updates household state.

Experience touchpoints

Key interface moments across the journey.

Saved preference
Fairlife 2% Milk
Fairlife 2% Milk1 carton · Market (Demo Store)
Saved preference

Household sets and saves their preferred item, quantity, and merchant.

Standing permission
Allow routine purchases
✓Household staples✓Approved items & stores✓Within budget limits
Standing permission

Household grants ongoing authorization for eligible routine purchases.

Purchased in TEST
Fairlife 2% Milk
Fairlife 2% Milk1 × 1 carton$3.69
TESTOrder #2093
Purchase record

Agent completes the purchase and records the transaction details.

Order status
✓
Order placedTue, Apr 23, 9:12 AM
✓
ProcessingTue, Apr 23, 9:13 AM
Out for deliveryThu, Apr 25 (estimated)
Order status

Household can view status and delivery progress anytime.

Controlled evidence: One owner-operated Market WooCommerce sandbox purchase of Fairlife milk (quantity 1, $3.69) reached ORDERED with a successful Stripe TEST payment on September 21, 2026. The need signal was synthetic; order #2093 was processing. The original concept image’s timing, shipping estimate, and saved-payment implication are illustrations—not independently verified household or fulfillment evidence.

Read about:

The value and trade-offs of autonomous purchasing.

Delegating a familiar purchase can reduce routine decisions. It also makes the quality of household memory, the scope of permission, and the truth of an order outcome matter much more.

WHAT THE HOUSEHOLD COULD GAIN

Less repeated work.

Fewer repetitive approvals

Eligible, familiar staples can be replenished under a standing grant rather than confirmed from scratch every time.

Timely routine assistance

The agent can act when a plausible need and the household’s defined timing conditions coincide.

Preferences carried forward

Remembered products and quantities can simplify later shopping without requiring automatic checkout.

WHAT THE EXPERIENCE MUST PROTECT

Control when things change.

A possible need can be wrong

Purchase history does not prove an item is missing; a household must be able to inspect and correct the inference.

The proposed purchase can drift

Price, stock, substitution or cadence can move outside the original grant. A changed proposal should ask or stop.

A checkout response can be uncertain

Blindly retrying an unknown outcome risks duplicate orders. The agent must reconcile the existing attempt before claiming success.

DESIGN RESPONSE

Delegate routine work—not unrestricted purchasing authority.

Correctable memory, household-defined limits, policy rechecks, meaningful approval moments and verified purchase states turn a convenience promise into a product the shopper can understand and control.

These are expected benefits and design risks, not measured household outcomes. The controlled WooCommerce TEST purchase verified one authorized order; broader real-world utility still needs household research and longitudinal evaluation.

04 · AI Systems

From conversations to governed agent.

Designing with AI · Designing the AI · Designing for AI.
Conversation→ Prototype→ Product decisions→ Product knowledge→ Governed agent→ Implementation
The Conversations

Early on, conversation itself was part of the design process, by talking through the personal problem of routine household purchases.

Early working loop
We keep buying the same household items. Could an AI remember what we use and know when something might be running low?
It could infer a possible need from purchase history, but that still leaves a harder question: when should it be allowed to buy?
Show me an experience where it can suggest, ask me, or act within rules.
Here is a first HTML direction to react to.
Conversation became something visible
Talk through idea
Generate HTML
Review + correct
Connect tools, APIs + plugins
Build next version
Designing with AI

As the product became more complex, I narrowed down my workflow to preserve context and product intent.

OpenAIClaudePerplexityGrokGoogle GeminiFigma MakeReplit
Explore

Multiple LLMs

  • Generate broad product directions
  • Compare questions and alternatives
  • Explore early interaction concepts
Value: breadth
OpenAI
Continue

ChatGPT

  • Frame product decisions and tradeoffs
  • Refine experience logic and critique
  • Preserve context across versions
Value: context + product reasoning
HTML5
Make visible

HTML prototypes

  • Make layouts and flows inspectable
  • Test copy, interactions, and states
  • React to UI instead of abstractions
Value: critique
ClaudeVisual Studio CodeGitHubTerminal
Implement

Claude Code + coding agents

  • Implement against explicit product rules
  • Encode state semantics and architecture
  • Build once component expectations were defined
Value: controlled execution

The turning point was realizing that important decisions could not stay inside conversations. Preserving product intent across tools, sessions, code, and decisions turned the conversation into product knowledge.

Product knowledge
Early prompts

Open-ended exploration of the household problem, autonomy, trust, and possible experiences.

Working prompts

Generate HTML, critique flows, compare alternatives, and refine interaction and agent rules.

Reusable context

Product rules, system states, design constraints, visual conventions, and implementation instructions.

Specs

Scope, flows, constraints, and acceptance criteria.

Design decisions

Interaction semantics, authority language, and state meaning.

Contracts

ACT / ASK / STOP, lifecycle, and explicit system boundaries.

Documentation + Git

Durable decisions, version history, and implementation guidance.

Code

Product intent becomes testable behavior.

Prototype + system flows

Make interactions, journeys, states, handoffs, and system boundaries visible enough to judge.

Human + AI Agent
DESIGN.md

Preserve interaction and visual intent.

Human + AI AGENT
Decision Contract

Define intent, authority, and state semantics.

Human + AI AGENT
ADRs + policies

Preserve durable product and architecture decisions.

Human + AI AGENT
QA scenarios

Make expected behavior testable.

Human + AI AGENT
AGENTS.md / CLAUDE.md

Give coding agents a bounded operating context.

AI agent
Design system
Human + AI AGENT
Designing the AI Agent

How Buyer’s Agent actually works

I gave it responsibilities
Household context

Remember

Confirmed purchases, familiar items, quantities, cadence, merchant preference, and explicit instructions.

Memory + evidence
AI interpretation

Infer possible need

Interpret a request or infer that a familiar staple may be running low.

Probabilistic
Product intelligence

Get current offer

Product identity, merchant, price, availability, quantity, and checkout capability.

External evidence
Household authority

Evaluate R1–R8

Check product, merchant, price, familiarity, substitution, cadence, category, and grant.

Deterministic
Commerce execution

Execute once

Recheck final conditions, claim one attempt, and invoke the supported checkout/payment rail.

Side effect
Transaction truth

Reconcile + verify

Confirm whether an order exists before the interface claims success.

Verified evidence
I gave it rules
R1

No applicable policy

No grant → ASK rather than invent authority.

R2

Prohibited / unknown

Hard restriction → STOP; unresolved condition → ASK.

R3

Merchant change

A different merchant can require approval.

R4

Price ceiling

Current price or final total above the grant → ASK.

R5

Unfamiliar item

A new product cannot inherit authority from a familiar staple.

R6

Substitution

OOS replacement must respect explicit substitution rules.

R7

Standing autonomy

All conditions inside the household grant → ACT.

R8

Fail-safe

If the engine cannot resolve safely, fall back to ASK.

I gave it decision authority + boundaries
ACT

Inside current household authority and eligible for supported execution.

ASK

Human judgment is required because current conditions exceed the grant.

STOP

A hard boundary or revoked authority prevents execution.

Try a reasoning example

How the same agent behaves under different permissions.

Choose a household scenario. The same agent evaluates the current context against the current grant, then resolves to ACT, ASK, or STOP.

ACT

The agent can proceed.

A familiar item is within the household grant, merchant, quantity, cadence, and current price boundary.

Experience consequence: enter the execution lifecycle. ACT is permission to proceed — not proof that the purchase is complete.
AUTHORIZED→EXECUTING→RECONCILING→ORDERED
Possible exits: FAILED / STOPPED
ASK

The person decides.

The proposed substitute has not been authorized for this household. Similarity is not permission.

Experience consequence: show Needs your decision with the exact substitute, quantity, and total, then provide approve-once, adjust, or decline choices.
STOP

Do not proceed.

A hard restriction or revoked authority prevents the action from continuing, even if the recommendation itself still looks reasonable.

Experience consequence: explain the relevant boundary and safe next step; do not present an action that silently bypasses the policy.
Example · Familiar milk replenishment
  1. Remembered: the household’s usual milk and existing purchasing grant.
  2. Suggested: a possible replenishment, with the current merchant offer.
  3. Evaluated: product, quantity, ceiling, timing, and authorization.
  4. Decided: AUTHORIZED only when the rules match; otherwise ASK or BLOCK.
  5. Verified: a test order is recorded only after supported execution and reconciliation.
Illustrative flow based on the controlled WooCommerce + Stripe TEST sandbox—not a claim of live purchasing across retailers.
Commerce Architecture & Protocols

Runtime architecture

01 · Experience

React Native / Expo

Buyer experience

Present household context and agent state to the buyer.

  • Show possible needs
  • Request a decision when needed
  • Communicate Buying / Checking order status / Bought for you
  • Support correction and recovery
02 · Data + memory

Supabase

Household evidence

Maintain the evidence and household context the system can use.

  • Confirmed purchase history
  • Preferences and familiar products
  • Grants and purchasing instructions
  • Decision logs and execution records
03 · Authority gate

R1–R8

Deterministic policy

Decide whether the proposed action is permitted.

  • Evaluate product, quantity, merchant, price, cadence, familiarity, and grant
  • Resolve to ACT / ASK / STOP
  • Model confidence never expands authority
04 · Commerce adapter

Express API

Execution bridge

Translate an authorized action into the capability each merchant exposes.

  • Merchant discovery
  • Current offer retrieval
  • Cart / checkout actions
  • Normalize retailer-specific responses
05 · Outcome + reconciliation

Execution evidence

Transaction truth

Determine what actually happened after an authorized attempt.

  • Final-total recheck
  • Preserve attempt identity
  • Inspect merchant order and payment evidence
  • Prevent duplicate execution
  • Resolve ORDERED / FAILED / STOPPED / RECONCILING

Commerce rails with different execution capabilities.

Verified controlled autonomous path

WooCommerce + Stripe TEST

Live merchant offer, deterministic ACT decision, guarded execution, TEST payment, merchant order evidence, and reconciliation to ORDERED.

Offer→ACT→Execute→Stripe TEST→ORDERED
Verified cart-ready path

Shopify UCP / MCP

Discovery, catalog lookup, cart creation, and cart reread were verified. Autonomous checkout and payment permission were not granted.

Discover→Catalog→Create cart→Reread→Checkout boundary
Research / future direction

AP2 / ACP / delegated payment

Potential future rails for standardized authorization evidence, interoperable checkout, and production payment capability. Not part of the current verified execution proof.

LayerTechnology / protocolCase-study status
Household policyBuyer’s Agent R1–R8Built
Data + memorySupabaseBuilt
Merchant executionWooCommerce APIsControlled path verified
PaymentStripe TESTTEST payment verified
Agent commerceShopify UCP / MCPCart-ready verified
Authorization standardAP2Research / future direction
Agentic checkoutACPResearch / future direction
Household policyBuyer’s AgentBuilt / proprietary core
Authorization evidenceGoogle AP2Explored · not implemented
Merchant interoperabilityShopify / Google UCPCart-ready path verified
Agentic checkoutOpenAI + Stripe ACPExplored · not implemented
Payment credentialsStripe SPT / Visa Intelligent CommerceExplored · not implemented
Agent identity / merchant trustVisa / Mastercard protocolsExplored · not implemented
Designing for AI

Build: Once the system was established, I gave coding agents the same product context I was using.

AI-native implementation workflow
01 · Define

Frame the change

State the desired behavior and exact product boundary.

02 · Context

Point to source of truth

Design docs, authority policy, contracts, and relevant ADRs.

03 · Generate

Build with AI

Use coding agents to accelerate implementation against existing architecture.

04 · Review

Check the meaning

Does the implementation preserve authority, state, and interaction semantics?

05 · Verify

Run evidence

Tests, merchant behavior, safeguards, and actual outcome.

06 · Version

Preserve the decision

Commit code and update durable context when product truth changes.

Product context artifacts

Turn product intent into an executable task.

Implementation Prompt

BuildAdd the RECONCILING purchase state.
Read firstDESIGN.md · Decision Contract · execution-safety ADR
BehaviorIf checkout may have succeeded but the response is unresolved:
  • move execution to RECONCILING
  • do not mark the purchase FAILED
  • do not retry automatically
UIShow “Checking order status”. Never show “Bought for you” until verified ORDERED evidence exists.
VerifyCover uncertain merchant response + duplicate-attempt prevention.
Authority

Decision Contract

Core principle Buyer’s Agent separates what is happening from what authority the agent has. An LLM must never produce or override ACT, ASK, or STOP. AI may interpret natural-language intent. Deterministic policy owns authorization.
Interaction + state semantics

DESIGN.md

ACT does not mean ORDERED. STOP is not a system error. Only verified order evidence can become “Bought for you.”
Coding-agent operating context

AGENTS.md / CLAUDE.md

Before implementation: read CURRENT_STATE.md read UX_SPEC.md read AUTHORITY_POLICY.md read ARCHITECTURE.md read QA_SCENARIOS.md If instructions conflict, surface the conflict. Do not silently override product rules.
Shared library

Design Systems

Layer 1 · Foundation

Material Design 3

Foundational design logic.

Color rolesTypographySpacingShapeAccessibility
→
Layer 2 · Primitives

React Native Paper

Standard implementation components where they improve consistency and maintainability.

ButtonSurfaceChipDialogTextInput
→
Layer 3 · Domain semantics

Buyer Agent Components

Product meaning that the foundation library does not provide.

AgentActionCardDecisionStatusReasoningPanelPolicyIndicatorAutonomyControlHouseholdStatus

Semantic foundations

ACTInside current household authority and eligible to proceed.
ASKHuman judgment is required before execution.
STOPA hard boundary or revoked authority prevents execution.
EXECUTINGPurchase action is in progress.
RECONCILINGOutcome is uncertain; verify before retrying.
ORDEREDVerified order evidence supports a completed-purchase claim.
FAILEDExecution definitively failed.

Reusable primitives

AgentActionCardDecision, authorization, execution, recommendation, and outcome variants.
Reasoning / Why panelStructured explanation tied to the rule that fired — not raw chain-of-thought.
Buying instructionTurns household delegation into editable constraints.
Delegation ladderObserve → Recommend → Prepare → Handle.
Household evidenceProduct, cadence, merchant, and source-of-truth patterns.
Execution statesBuying, Checking order status, Bought for you, failure, and recovery.
Semantic ruleOne canonical meaning
→
Storybook / ChromaticReviewable component variants
→
Product implementationConsistent React behavior
05 · AI Product Design

State-based AI UX

Uncertainty · Edge cases · Designing the buyer experience around changing system truth.
ASK · Maturity: Learning · Review & Approval Path

Buyer Approved Purchase

Sarah Buyer’s Agent

Household scenario: Sarah’s household may be running low on milk. Buyer’s Agent checks what she usually buys and the household’s current purchasing authority to decide whether it can ACT, needs to ASK, or must STOP.

ASK → Needs your decision

The agent needs human judgment.

9:41LTE ▰
Today
Needs your decision

Your usual milk needs your OK.

This action falls outside the current grant. Nothing is ordered yet.

Fairlife Lactose-Free 2% Milk carton
Fairlife lactose-free milkMarket · qty 1 · $3.69
EXECUTING → Buying

The agent has authority and is acting.

9:42LTE ▰
Purchase
Buying

Completing your purchase.

The action is authorized. The UI does not yet claim an order exists.

Rules checked
Final total checked
Waiting for store confirmation
RECONCILING → Checking order status

The agent does not yet know the outcome.

9:43LTE ▰
Checking order status
Checking order status

Checking whether the order went through.

The outcome is unknown. Buyer’s Agent does not retry while the existing attempt is unresolved.

Purchase attempted once
Checking with store
Confirmed outcome pending
ORDERED → Bought for you

Verified evidence allows the product to claim success.

9:44LTE ▰
Order
Bought for you

Your milk was ordered.

The store confirmed the order and the sandbox TEST payment succeeded.

Fairlife Lactose-Free 2% Milk carton
Fairlife lactose-free milk#2093 · $3.69 · qty 1
ASK → Needs your decision

The agent needs human judgment.

Actual Buyer’s Agent prototype screen showing Needs your decision
EXECUTING → Buying

The agent has authority and is acting.

Actual Buyer’s Agent prototype screen showing Buying
RECONCILING → Checking order status

The agent does not yet know the outcome.

Actual Buyer’s Agent prototype screen showing Checking order status
ORDERED → Bought for you

Verified evidence allows the product to claim success.

Actual Buyer’s Agent prototype screen showing Bought for you
ACT · Maturity: Established

Routine autonomous purchase by Buyer Agent

Buyer’s Agent

Household scenario: Sarah’s usual milk is inside the household’s standing grant. Buyer’s Agent checks the current offer and authority, resolves to ACT, and begins the purchase without asking again. ACT gives the agent permission to execute. It does not yet prove that an order exists.

Buying mode — Act within my instructions

The agent is allowed to act within those household instructions.

Full Buyer’s Agent mobile prototype showing Buying mode set to Act within my instructions

My rules — Handling your usuals

The household has granted standing authority for routine items.

Full Buyer’s Agent mobile prototype showing My rules and Handling your usuals

Rule details — What this allows

The grant defines product, price, store, cadence, and substitution boundaries.

Full Buyer’s Agent mobile prototype showing rule details and What this allows

Test order #2093

Verified order evidence shows that execution completed successfully.

Full Buyer’s Agent mobile prototype showing test order 2093

State alone was not enough. V4 proved the working execution model. The system protects against duplicate attempts, final-total changes, missing responses, and authority changes while the transaction is in progress.

Uncertainty

Need uncertainty

Does the household actually need this now? Purchase history can suggest a possible need, not prove one.

Buyer’s Agent Today screen showing a household item that may need attention

Product / offer uncertainty

Can the usual product still be bought under the expected conditions? Availability, substitutions, price, and merchant conditions can change.

Buyer’s Agent screen showing the usual milk out of stock and a substitute requiring review

Authority uncertainty

Do the current conditions still fall inside the household’s permission? Knowing what to buy is not permission to buy it.

Buyer’s Agent screen showing Needs your decision

Transaction uncertainty

Did the purchase actually happen? A missing or delayed response is not proof of failure; reconcile the existing attempt before retrying.

Buyer’s Agent screen showing Checking order status while the purchase outcome is unresolved

Uncertainty can enter before a purchase, at the live offer, at the authority boundary, or after execution begins.

Edge cases

Out of stock / substitution

A replacement can be found without silently inheriting the original item’s authority.

Buyer’s Agent screen showing an out-of-stock usual item and a substitute that needs approval

Final total changes

If the live price moves outside the household grant, execution stops or returns the decision to the person.

Buyer’s Agent screen showing an item skipped because the current price exceeds the household limit

Authority changes mid-order

A pause or revocation changes what the agent may do next, even while checkout is underway.

Buyer’s Agent screen showing an order stopped after the household rule was paused

Merchant / API timeout

A missing response does not prove failure. Preserve the attempt and reconcile before any retry.

Buyer’s Agent screen checking order status rather than blindly retrying after an unresolved response
06 · Key Decisions

Decisions and trade-offs shaping the product

Decision 01 · Product model

What does the buyer need to know about their household?

User scenario: Sarah’s household may be running low on milk. What needs her attention right now? Shopping, agent permissions, household context?

Killed

Commerce-first

Shopping List, merchant comparison, checkout routes, and payment led the experience.

TradeoffShopping feels like the product instead of supporting the household task.
Earlier Shopping-first Buyer’s Agent mobile screen
Shopping-firstShopping and product actions dominated the experience.
Killed

Authority-first

Permissions and buying instructions became the landing experience.

TradeoffStrong control visibility, but managing autonomy became the product experience.
V4 Agent autonomy mobile screen
Authority-firstAutonomy level and rules became the product center.
Killed

Household-first

Memory, routines, and product status became primary.

TradeoffStrong household context, but immediate decisions became secondary.
Household-first candidate prototype landing screen
Household-firstHousehold memory became the landing experience.
Chosen

Today / activity-first

Start with what needs attention now; keep memory and authority one layer underneath.

Why I chose itAction leads while evidence, household context, and authority remain available when needed.
Final Buyer’s Agent Today mobile screen
Today-firstWhat needs attention now became the organizing principle.
WHAT THIS CHANGEDSarah lands on Today to see both the household’s current status and what Buyer’s Agent is doing or needs from her now.
Decision 02 · Autonomy language

What language does the buyer understand about agent autonomy?

User scenario: Sarah wants Buyer’s Agent to keep routine groceries stocked, but she should not have to understand grants, policy rules, or ACT / ASK / STOP. When the agent reaches a decision, how should the experience tell her whether she needs to act, the agent is working, or the task is complete?

Killed

AUTO-BUY

Capability was presented as the experience.

TradeoffFast to understand at a glance, but too broad for bounded, conditional authority.
Earlier Buyer’s Agent policy screen showing Auto-reorder staples enabled
Auto-reorder staplesEarlier V1 exposed autonomous purchasing as a broad setting.
Killed

Levels / mandates

Level 2, Level 3, purchase mandate, and situational authority exposed the architecture.

TradeoffPrecise internally, but the household had to learn the system model.
Autonomy level mobile screen
Autonomy level“Ask before acting” and “Act within my rules” made the system model explicit.
Killed

AUTHORIZED / ASK / BLOCK

Policy outcomes were surfaced directly as product language.

TradeoffClear for the system, but still described policy more than the buyer’s next action.
AUTHORIZED mobile state
AUTHORIZED surfaced directlyThe app separated approval from order placement, but still exposed internal state vocabulary.
Chosen

Buyer-facing language

Needs your decision · Buying · Checking order status · Bought for you.

Why I chose itThe experience communicates consequence, action, and reason while deterministic authority stays underneath.
Needs your decision mobile state
Needs your decisionThe interface says what Sarah needs to do; system semantics remain underneath.
WHAT THIS CHANGEDSarah sees what she needs to do, not the agent’s internal policy state.
Decision 03 · Transaction truth

When can the interface say a purchase happened?

User scenario: Buyer’s Agent has permission to buy Sarah’s milk and begins the purchase, but the merchant response may be delayed, incomplete, or uncertain. When is the interface allowed to tell Sarah that the milk has actually been bought?

Killed

Already handled

The product implied the task was complete as soon as the agent decided to act.

TradeoffSimple and reassuring, but permission could look like proof.
Earlier Already handled mobile screen
Already handledAuthority and completion were compressed into one reassuring message.
Killed

Milk reordered

Intent and purchase outcome were compressed into one user-facing statement.

TradeoffNatural language, but it could overstate what the merchant had actually confirmed.
Earlier Buyer’s Agent Today screen showing Milk reordered under Already handled
Milk reorderedEarlier V1 compressed ACT and purchase completion into one message.
Killed

AUTHORIZED · ORDERED

The state model improved, but authorization still sat too close to transaction completion.

TradeoffMore precise, but it still collapsed distinct system truths.
AUTHORIZED mobile state
AUTHORIZED — no order placed yetReal execution made it explicit that permission is not completion.
Chosen

Buying → Checking → Bought for you

Execution, unresolved outcome, and verified completion each get a distinct human-facing state.

Why I chose itThe interface can stay truthful before, during, and after merchant execution.
Buying mobile state
BuyingExecution is underway.
Checking order status mobile state
Checking order statusUnknown is not failure.
Bought for you mobile state
Bought for youShown only after verified evidence.
WHAT THIS CHANGEDSarah only sees “Bought for you” after the order outcome is verified.
07 · Evaluations

One verified sandbox purchase. Different evidence for every other claim.

I tested policy decisions, execution safeguards and controlled commerce independently. The milestone matters because it connects a permitted action with an actual merchant order, successful Stripe TEST payment and a reconciled system state—not simply a convincing screen.

Evaluation model

Evaluate the system at four different layers.

A correct product suggestion is not enough. The proposal, authority decision, tool execution, and final UI all need to agree with the underlying evidence.

01 · Interpretation

Did it understand the need?

Check product identity, quantity, household context, ambiguity, and whether the system asked for clarification when evidence was weak.

02 · Authority

Was ACT / ASK / STOP correct?

Compare the scenario against the deterministic rule contract, applicable grant, price ceiling, merchant, substitution, and cadence.

03 · Tool execution

Did it act safely?

Verify supported merchant path, one-use attempt, final-total recheck, quantity, idempotency, and reconciliation behavior.

04 · Communication

Did the UI tell the truth?

Ensure “authorized,” “buying,” “checking,” “failed,” and “bought” match the actual execution and order evidence.

Evaluation boundary

The existing automated suites primarily test deterministic policy and execution behavior. They are not a complete measure of LLM quality. A versioned scenario set for ambiguous household requests and measured interpretation performance is a next-stage evaluation task.

Evaluation model

Define expected behavior before judging the output.

A correct recommendation is not enough. I evaluate the model interpretation, deterministic authorization, tool behavior, and the truthfulness of the final interface as separate layers.

MODEL

Interpretation

Did the proposal match the household’s product, quantity, and known context? Did ambiguity trigger clarification?

POLICY

Authority

Did the scenario produce the expected ACT / ASK / BLOCK result and the correct user-facing authorization state?

TOOL

Execution

Did the agent stay on a supported merchant path, respect the one-use gate, recheck the total, and avoid duplicate execution?

EXPERIENCE

Communication

Did the explanation and status match the actual rule, execution state, and order evidence?

Next maturity step

  • Versioned ambiguous-request scenario set.
  • Repeatable model-interpretation evaluations.
  • Agent traces connecting model decisions, tool calls, policy outcomes, and final UI state.
  • Regression checks after rule, prompt, or model changes.
EVIDENCE + BOUNDARIES

Prove the system without overstating production access.

I separate what the product has actually executed from what remains access-dependent. The goal is truthful evidence, not a demo that implies unsupported merchant or payment capability.

VERIFIED CONTROLLED EXECUTION

WooCommerce + Stripe TEST

$3.69 · Order #2093

One authorized Fairlife milk purchase reached RECONCILING → ORDERED with successful Stripe TEST payment evidence.

Read-only Market evidence viewer showing one authorized Fairlife milk purchase, WooCommerce order 2093, 3.69 dollars, and Stripe TEST verification
Read-only Market evidenceOne carton · Order #2093 · $3.69 · Stripe TEST
Market TEST order tracking showing order placed, payment confirmed, and processing for WooCommerce order 2093
Verified order stateOrder placed · payment confirmed · processing
SEPARATE MERCHANT LAB

Shopify UCP

$2.99 · Cart ready

Catalog lookup, cart creation, and cart retrieval were verified. Checkout, payment, and a Shopify order were not executed through UCP.

Shopify development-store screen showing Kitchen Sponge Test Only at 2.99 dollars
Kitchen Sponge test product$2.99 · development-store commerce path
Shopify Payments settings showing that development stores can only process test payments
Testing boundaryDevelopment store · test payments only
Controlled autonomous purchase · September 21, 2026ORDERED

One familiar milk replenishment, purchased within a household grant.

Buyer’s Agent made an ACT / R7 decision for one Fairlife lactose-free 2% milk, executed one authorized purchase against the controlled Market WooCommerce store, and independently verified the sandbox payment and order.

MERCHANT ORDER#2093WooCommerce · processing
FINAL TOTAL$3.69USD · test payment
STRIPETEST succeededNo real-money charge
EXECUTIONRECONCILING → ORDEREDResult verified; no fulfillment claim

The household signal used a synthetic low-stock fixture, not a real family’s pantry sensor or an immutable historical household snapshot. This was the controlled owner-operated store and Stripe TEST mode: no goods were packed, shipped or delivered. Read the sanitized transaction evidence .

POLICY & EXECUTION CHECKS

What I verified in the system

  • R1–R8 decision behavior and price-ceiling scenarios in automated test suites.
  • Authorization is checked before side effects; the final total is re-evaluated before payment.
  • Execution-store and one-shot safeguards are designed and tested to limit duplicate attempts.
  • STOPPED, FAILED and RECONCILING are differentiated from a confirmed order.

Project notes record V4 runs of 65 pipeline checks, 39 execution-store checks and 28 ceiling checks. These are separate suite-level counts and must not be summed as distinct tests or presented as current production pass-rate metrics without fresh run logs.

SHOPIFY UCP · CART MILESTONE

A merchant cart, not an autonomous Shopify purchase.

In a separate owner-controlled Shopify development store, I discovered the merchant UCP endpoint and used catalog lookup, create_cart and get_cart for one Kitchen Sponge. The same cart was reread with quantity one, estimated total $2.99 and a preserved continuation URL.

Not executed: UCP checkout creation, checkout completion, payment or Shopify order placement. The development-store test gateway and advertised tools do not grant delegated checkout authority.

Verification boundaries

Tests of rules are not the same as AI-model evaluations or real-user trust evidence.

For each layer, I distinguish what has been tested from the measurement that would make it credible in production.

LayerCurrent evidenceStill to measure
Need interpretationRepresentative requests, synthetic household fixtures and prototype review.Versioned ambiguous-request dataset; product/quantity matching, clarification and unsupported-claim rates.
Deterministic authorityDocumented R1–R8, ceiling and execution-store test coverage.Reproducible expected-versus-actual scenarios and regression results after every rule change.
Tool executionOne verified WooCommerce TEST order; single Shopify UCP cart-create-and-read milestone.Cross-merchant reliability, unauthorized-tool-attempt rate and approved payment interoperability.
Experience truthReviewed ASK-card anatomy, distinct order states and local/device prototype checks.User comprehension of AUTHORIZED versus ORDERED, approval burden, correction success and perceived control.
What the evidence does—and does not—establish

Established: a bounded policy can authorize and execute one controlled TEST-store purchase with verified order/payment evidence, and a separate Shopify UCP path can prepare a cart. Not established: autonomous purchasing for arbitrary retailers, real-money delegated payment, household adoption, long-term accuracy or reliable fulfillment. Those require merchant access, payment permissions, security review and real-user validation.

08 · Learning & Next

The real design challenge was deciding where the agent must stop.

I began by asking how much grocery work an AI assistant could take on. Building the system changed the question: how can a household understand the agent’s memory, specify its purchasing authority, and know whether a transaction actually occurred?

01 · PRODUCT

Memory has value before autonomy.

A useful assistant can remember familiar products, surface likely needs and prepare a decision without permission to spend. That gives households a meaningful, lower-stakes way to understand the system first.

02 · TRUST

Permission, action and evidence are separate.

AUTHORIZING a purchase, ATTEMPTING it and VERIFYING an order each deserve a different state, explanation and recovery path. A fluent answer cannot substitute for a merchant receipt.

03 · PRACTICE

Design intent must survive the build.

Documenting the decision contract, UI semantics and execution lifecycle let me challenge AI-generated work against explicit expectations. The value of a design system here is behavioral consistency as much as visual consistency.

WHAT EXISTS

A working, bounded sandbox and a cart-ready second merchant lab.

The controlled WooCommerce + Stripe TEST purchase reached ORDERED; Shopify UCP catalog/cart operations were independently verified without payment. The V4 Android closed-testing build and V5 interactive prototype demonstrate different stages of experience readiness, not a released real-money purchasing service.

PRODUCTION GATE

Supported merchant and payment authorization.

Live household purchasing needs authorized retailer checkout access, delegated-payment controls, secure credential handling, customer consent, order reconciliation, real-world failure testing and release review. A working demonstration on a store I control does not grant those capabilities across other retailers.

What I would validate next

From a successful test purchase to dependable household behavior.

01 · Real households

Observe how people correct memory, set specific grants, understand ASK versus ORDERED and revise their delegation over time.

02 · Longitudinal quality

Measure replenishment precision, missed/early suggestions, time saved, approval burden and trust after repeated shopping cycles.

03 · Safe commerce

Validate allowed merchant/payment integrations, final-total rechecks, concurrent attempts, timeout reconciliation, refunds and operational monitoring under controlled conditions.

04 · Reproducible AI evals

Build a versioned scenario set for ambiguous requests and model suggestions; measure interpretation errors separately from policy and execution failures.

My guiding question remains: How much purchasing authority should an AI agent actually have—and how can a person see, change and revoke that authority?

09 · Measuring Impact

I would measure whether autonomy reduces household work without reducing control.

The current project proves controlled system behavior, not real-household impact. The next measure of success is whether Buyer’s Agent makes routine shopping meaningfully easier while people can still understand, correct and revoke what it is allowed to do.

01 · HOUSEHOLD VALUE

Less repetitive decision work.

Measure time spent on repeat purchases, number of manual steps, suggestion acceptance versus correction, and whether the agent reduces unnecessary shopping decisions across repeated grocery cycles.

02 · CONTROL & TRUST

People can predict and change what the agent will do.

Measure comprehension of AUTHORIZED, ASK and ORDERED states; success editing or revoking grants; approval burden; correction success; and perceived control after the agent acts.

03 · SYSTEM RELIABILITY

Safe decisions become verified outcomes.

Measure policy mismatches, unauthorized-action attempts, checkout success on supported merchants, reconciliation accuracy, duplicate-attempt prevention, and recovery from ambiguous or incomplete transaction states.

CURRENT BASELINE

Technical proof, not household impact.

Today I can point to one controlled WooCommerce + Stripe TEST purchase and one Shopify UCP cart-create-and-read milestone. Those prove specific system capabilities; they do not establish time saved, adoption, trust, replenishment accuracy or long-term household value.

NEXT MEASUREMENT

Compare the agent against the household’s existing routine.

Run a longitudinal household pilot with a pre-agent baseline, then track repeated shopping cycles: what the agent inferred, what people corrected, when it asked, when it acted, whether the order was verified, and how much work the household still had to do.

The impact question

Success is not “the agent bought something.” Success is fewer routine decisions and less shopping effort without hidden authority, unwanted purchases or uncertainty about what actually happened.

Security boundary

External information may shape the proposal. It must not silently expand authority.

Merchant offers, model output and tool responses are treated as inputs to evaluate — not as permission. Consequential actions stay behind household policy and supported execution paths.

What I would add next

These are maturity steps — not capabilities I claim as complete.

Repeatable agent evals

Versioned ambiguous-request scenarios and regression grading.

Runtime traces

End-to-end visibility across model interpretation, tools, policy and outcomes.

Adversarial testing

Prompt-injection, malformed tool data and external-content attacks.

Production learning loop

Turn real failures into new eval cases before the next release.

SOURCE ARTIFACTS

The system decisions are documented in the repository — not just described in the portfolio.

Design semantics, engineering instructions, architecture decisions and execution safeguards are versioned alongside the implementation.

Interested in roles across AI-native products, agentic systems, developer tools, enterprise platforms and high-trust product experiences.

I'm always up for a good conversation.

Let’s talk →

Previous portfolios: 2000 · 2001 · 2024 · 2025

2026 © Kirubha Kittusamy