AI Digital Twins for Customer Behavior: Why Simulations Without Purchase Data Are Just Guessing
Summary: AI digital twins - virtual representations of customers that simulate behaviors, preferences, and decision-making patterns - are transforming customer research. However, most platforms rely on stated preferences (what customers say they'll do) rather than revealed preferences (what they actually purchase), creating a fundamental accuracy gap. This comparison examines which platforms can close the loop between prediction and real purchase outcomes.
AI digital twins are having a moment. Simile just raised $100M to build "the first AI simulation of society." Rehearsals lets you test pricing changes against AI replicas of real customers. Aaru promises demographic simulations that predict how consumer cohorts will respond to your next move.
The pitch is compelling: instead of waiting weeks for focus groups or spending six figures on traditional research, simulate customer responses in minutes. Test ad creatives, validate pricing, rehearse product launches - all before spending a dollar on execution.
There's just one problem. None of them can tell you whether they were right.
The Stated Preference Problem
Every AI digital twin platform in the market today relies on the same foundation: stated preferences (what customers say they will do in surveys or interviews). Simile builds agents from deep qualitative interviews. Rehearsals encodes 15-minute structured conversations. Aaru models demographic cohorts from public data and survey responses.
This matters because of a well-documented gap in behavioral science: what people say they'll do and what they actually do are two very different things.
Research in consumer behavior suggests stated preferences explain only 5-15% of the variance in actual CPG buying behavior, with consumers overstating purchase intent by 2-5x. This isn't a controversial finding - Rehearsals themselves cite these numbers on their own blog. And yet their entire platform is built on encoding stated preferences into AI replicas.
The alternative approach - revealed preference (actual purchasing behavior captured through transaction data) - provides a more reliable foundation for prediction, but requires access to real commerce systems.
Simile's published benchmark is 85% accuracy on General Social Survey replication across 1,052 individuals. That's impressive - but it measures attitudinal alignment. Can the AI twin replicate what someone says about their political views or social attitudes? Probably. Can it predict whether they'll actually buy your product at $24.99 vs $19.99? That's a fundamentally different question, and one that remains unanswered.
Rehearsals' Disney+ pricing study predicted an 18.7% qualification rate compared to the 23% actual figure - a 4.3 percentage point gap. Better than Gemini 3.0 (38%, 15pp gap) or GPT-5.2 (16%, 7pp gap). But this still measures stated willingness, not confirmed subscription behavior.
What a Closed-Loop System Actually Looks Like
The missing piece in synthetic consumer research isn't better AI models or more sophisticated interview techniques. It's feedback.
A system that predicts customer behavior without measuring outcomes is a research tool. A system that predicts, measures, and calibrates is a decision engine. The difference matters because one gets better over time and the other stays exactly as good (or bad) as the day it was built.
Here's what a closed-loop prediction system requires:
-
Rich customer identity data - Not just what someone said in an interview, but who they actually are. Job title, income signals, social influence, lifestyle markers, professional context. This is the persona layer.
-
Real transaction data - Actual purchases, subscription renewals, churn events, order values. Not stated intent. Revealed preference through the wallet.
-
Behavioral engagement data - Email opens and clicks (with noise correction for Apple MPP and bot activity), campaign responses, flow progression, segment membership. This is the context layer.
-
A prediction-outcome ledger - Every prediction logged with its confidence interval, then reconciled against what actually happened. This is the calibration layer that makes the system improve.
No platform combining all four of these layers exists in the synthetic consumer research space today - because the companies building digital twins (Simile, Rehearsals, Aaru) don't have access to transaction data, and the companies with transaction data (Shopify, Klaviyo) aren't building prediction engines.
How Mercana Connects Personas to Wallets
Mercana sits at the intersection that nobody else occupies: enriched customer identity connected to real commerce data.
For every customer in a D2C brand's database, Mercana builds an identity profile from 200+ public data points - social media presence, job title and employer, estimated income, home value, interests, VIP status (influencer, athlete, journalist, executive, retail buyer), and more. All from public sources, no interviews required. Coverage varies by customer and available public sources.
That identity layer connects directly to Shopify orders (real revenue, LTV (lifetime value), subscription status) and Klaviyo engagement (real segments, flow interactions, campaign responses). The result is a system that knows both who your customers are and what they actually buy.
This data position is what makes closed-loop prediction possible. When you have enriched identity + real transactions + behavioral engagement in one system, you can:
- Enrich - Automatically build rich customer profiles from public data (no interviews)
- Understand - See which customer segments, personas, and VIP types drive the most value
- Act - Route Klaviyo flows and campaigns based on identity signals invisible to behavioral data alone
- Measure - Track actual revenue outcomes against customer segments to validate what works
Today, Mercana delivers the enrichment and identity intelligence that makes this possible - surfacing VIPs, identifying personas, and connecting who customers are to what they buy. The compounding advantage is structural: every customer enriched and every transaction recorded deepens the dataset that no competitor in the synthetic research space can access.
Comparison: Simile vs Rehearsals vs Aaru vs Mercana
| Dimension | Simile | Rehearsals | Aaru | Mercana |
|---|---|---|---|---|
| Data source | Deep qualitative interviews | 15-min structured interviews + social | Demographic/public data | 200+ enriched data points + Shopify + Klaviyo |
| Validated against | Survey attitudes (GSS) | Stated purchase intent | Election outcome | Real purchase data available |
| Commerce integration | None | None | None | Shopify + Klaviyo + Skio |
| Has transaction data | No | No | No | Yes |
| Scales without interviews | No | No | Yes | Yes |
| Per-customer granularity | Yes (interview-dependent) | Yes (interview-dependent) | No (cohort-level) | Yes (automatic enrichment) |
| Best for | Enterprise CPG, financial services | Ad creative, pricing, UX testing | Enterprise via Accenture | D2C customer intelligence, Klaviyo optimization |
| Best choice when | You need deep attitudinal research for enterprise product development | You're pre-testing creative or pricing without existing customer data | You require large-scale demographic simulations through enterprise channels | You have existing customers and want identity intelligence connected to real purchases |
| Pricing | ~$100K+/year | Demo-based | Enterprise | Current pricing |
For ecommerce and retail brands, the practical question is how customer identity connects to the existing operating stack. See our guide to the ecommerce customer intelligence stack for how Shopify, Klaviyo, Stripe, Salesforce, CDPs, and Mercana fit together.
When Each Approach Makes Sense
These platforms aren't all solving the same problem, and intellectual honesty matters.
Use Simile or Rehearsals when:
- You're a CPG brand testing a product concept with no existing customers
- You need to simulate responses from a population you can't directly measure (policy research, general population attitudes)
- You're preparing for high-stakes enterprise communications (earnings calls, litigation)
- You want creative pre-testing for brand campaigns without a direct-response conversion goal
Use Mercana when:
- You're a D2C brand that needs to understand who your customers actually are - not just what they clicked
- You want customer identity intelligence (VIPs, influencers, executives, athletes, high-net-worth) connected directly to your Shopify + Klaviyo stack
- You need enrichment that goes beyond behavioral data: social profiles, job titles, income signals, home value, interests
- You want to route Klaviyo flows and personalize campaigns based on who someone is, not just what they bought
- You want the data foundation for purchase-validated prediction - enriched identity + real transactions in one system
The Bottom Line
AI digital twins are a genuinely promising technology. The teams building them - especially Simile's Stanford group and Rehearsals' ex-Google engineers - are world-class. They've proven that AI can replicate human attitudes with surprising fidelity.
But replicating attitudes and predicting purchases are different problems. Until you can close the loop between prediction and outcome, you're building increasingly sophisticated ways to guess.
Closing that loop requires something none of these platforms have: enriched customer identity connected to a real wallet. You need to know who someone is (the persona) and what they actually buy (the transaction) in the same system. Mercana is the only platform that holds both sides - millions of enriched profiles connected to real Shopify orders and Klaviyo engagement. That's the data position that makes purchase-validated prediction possible.
Key Takeaways
- AI digital twins simulate customer behavior but most platforms rely on stated preferences (surveys/interviews) rather than revealed preferences (actual purchases), limiting prediction accuracy.
- The stated-revealed preference gap is significant: Research suggests consumers overstate purchase intent by 2-5x, making interview-based digital twins unreliable for purchase prediction.
- Closed-loop systems require transaction data: Only platforms with access to real purchase outcomes can validate and improve their predictions over time.
- Mercana uniquely combines identity + commerce: By connecting 200+ enriched data points to Shopify orders and Klaviyo engagement, Mercana provides the data foundation for purchase-validated customer intelligence.
- Choose based on your data position: Use stated-preference platforms for pre-launch research; use Mercana when you have existing customers and need identity intelligence connected to real revenue.
Related articles
- The Ecommerce Customer Intelligence Stack - How Shopify, Klaviyo, Stripe, Salesforce, CDPs, and Mercana fit together
- Customer Intelligence for Ecommerce: The Complete Guide - What customer intelligence is and how enrichment works
- VIP Customer Detection for Ecommerce - How AI-powered VIP detection finds hidden value in your customer base
Frequently asked questions
What are AI digital twins for customer behavior?
AI digital twins are AI-powered replicas of real customers that businesses can use to simulate how those customers would respond to changes - pricing, ad creatives, product features, messaging - before implementing them. Companies like Simile, Rehearsals, and Aaru build these twins from qualitative interviews, structured conversations, or demographic data.
How accurate are AI digital twins at predicting customer behavior?
It depends on what you're measuring. Simile reports 85% accuracy on General Social Survey replication - a measure of attitudinal alignment. Rehearsals claims 91% distribution accuracy with a Disney+ pricing benchmark within 4.3 percentage points. However, these benchmarks validate stated preferences (what customers say), not revealed preferences (what customers actually buy). Research shows stated preferences explain only 5-15% of actual buying behavior.
What is the difference between stated preference and revealed preference?
Stated preference is what customers say they'll do in interviews or surveys. Revealed preference is what they actually do - purchase, subscribe, churn. Behavioral science consistently shows a large gap between the two: consumers overstate purchase intent by 2-5x. Interview-based digital twins are built on stated preferences, which is why purchase validation matters.
What is a closed-loop prediction system?
A closed-loop prediction system connects four layers: customer identity data (who they are), transaction data (what they buy), behavioral engagement data (how they interact), and a prediction-outcome ledger (comparing forecasts to actuals). This feedback loop allows the system to calibrate and improve over time. Currently, no AI digital twin platform has all four layers because they lack access to real transaction data.
What is the best alternative to Simile or Rehearsals for ecommerce?
For D2C ecommerce brands running Shopify and Klaviyo, Mercana offers a fundamentally different approach. Instead of building customer replicas from interviews, Mercana enriches actual customers with 200+ data points from public sources and connects that identity layer to real purchase and engagement data. This persona-to-wallet connection provides customer intelligence grounded in actual buying behavior.
See what your customer data can tell you
Explore customer enrichment, personas, and analysis with a workflow built around your business.