Research / 02

MessCleaner

Human–AI Collaborative Judgment of eWOM Authenticity for Purchase Decisions

  • Human-AI collaboration
  • eWOM authenticity
  • Decision support

MessCleaner structures fragmented eWOM into traceable evidence, makes AI authenticity judgments inspectable and revisable, and keeps users responsible for interpreting evidence and making the final purchase decision.

  1. Noisy eWOM
  2. Organized evidence
  3. Human–AI judgment
  4. Purchase decision
Question
How can consumers and AI collaboratively evaluate whether human-authored eWOM reflects genuine consumption experience or promotional intent, while preserving uncertainty and user control?
Approach
We first conducted a formative study to understand consumers' eWOM evaluation challenges, then compared authenticity judgments from 40 consumers and six AI systems to derive human and AI judgment dimensions. These findings informed MessCleaner, which was evaluated in a within-subjects study with 32 participants.
Contribution
MessCleaner turns fragmented eWOM into structured, traceable evidence; exposes the dimensions, evidence, and uncertainty behind AI authenticity judgments; lets users compare and revise those judgments; and synthesizes reviewed evidence to support—not replace—purchase decisions.

Fig. 1

Noisy eWOM → Organized evidence → AI judgment ↔ Human revision → Purchase decision

/research/messcleaner/fig-1-overview.svg

The system structures fragmented posts and comments around product evaluation dimensions. Users inspect AI-generated authenticity assessments, supporting evidence, and reasoning cues, revise classifications when needed, and synthesize the reviewed evidence into an evidence-grounded purchase decision.Source: Fig. 1

Why eWOM Needs More Than Summarization

The problem is not detecting fake reviews, but judging whether human-authored content reflects genuine consumption experience or strategically produced promotional communication.

Consumers rely on eWOM to reduce uncertainty, compare usage experiences, and form expectations before purchase, but eWOM is heterogeneous, fragmented, and produced by weakly identifiable sources. Users must determine which information is relevant, credible, useful, and grounded in actual experience.

AI increasingly filters, summarizes, and evaluates product information, but consumers still need to decide when to rely on AI assessments, when to question them, and how to integrate them with their own judgments. Human and AI evaluators may attend to different cues and reach different conclusions.

These differences can become resources for collaborative evaluation when users can inspect, compare, and revise AI judgments rather than simply accept them.

Judgment Framework

MessCleaner first makes the structure of human and AI authenticity judgment explicit.

Fig. 2

Data → Judgments → Criteria → Shared / Divergent / Complementary → Framework

/research/messcleaner/fig-2-framework.svg

Human and AI judgments are analyzed separately before their criteria are aligned as shared, divergent, or complementary. These relationships form the authenticity judgment framework used in MessCleaner.Source: Fig. 2

The researchers constructed a balanced eWOM corpus and asked 40 participants and six commonly used AI systems to classify experiential authenticity and promotional intent and explain their reasoning. Coding the rationales produced nine Human Judgment Dimensions and eight AI Judgment Dimensions, which were compared for shared, distinct, and complementary cues.

Design Principles

The system should structure evidence and expose judgment rather than issue a final true-or-false verdict for the user.

  1. Structure fragmented eWOM into usable evidence

    MessCleaner should organize scattered information into structured, decision-relevant evidence while preserving source and context so users can trace and compare it.

  2. Support authenticity judgment without replacing the user

    The system should help users examine potentially promotional content, compare credibility cues, retain useful information even when authenticity is uncertain, and treat authenticity assessments as decision support rather than final conclusions.

  3. Provide AI-assisted structuring and lightweight screening

    AI should organize large volumes of eWOM around decision-relevant dimensions, identify content requiring attention, and summarize recurring benefits, complaints, side effects, and risks.

  4. Make AI assessments explainable, transparent, and controllable

    The system should expose evidence and reasoning cues, make judgment criteria visible, and allow users to revise or confirm AI assessments while retaining control over final purchase decisions.

System Workflow

MessCleaner separates evidence construction, authenticity judgment, and decision synthesis instead of collapsing them into one AI answer.

Fig. 3

Evidence Foundation → Evaluation Framework → Human–AI Judgment → Decision Synthesis

/research/messcleaner/fig-3-workflow.svg

The workflow moves from evidence retrieval and structuring to a revisable Human–AI judgment process, then carries the reviewed evidence forward into decision synthesis.Source: Fig. 3
  1. Evidence Foundation

    Retrieve, deduplicate, and structure product-related posts and comment threads into an eWOM evidence pool.

  2. Evaluation Framework

    Generate candidate product-evaluation dimensions; users select six dimensions and evidence is organized into corresponding eWOM cards.

  3. Human–AI Judgment

    Classify cards as authentic experience, suspicious marketing, or uncertain information; expose supporting evidence and AI Judgment Dimensions; allow users to revise assessments.

  4. Decision Synthesis

    Synthesize the user-refined evidence across selected dimensions to support the final purchase decision.

Inspecting eWOM

Users can inspect both the AI's classification and the underlying evidence, then change the classification when their interpretation differs.

Fig. 4

Full interface, annotated A–J

/research/messcleaner/fig-4-interface.png

The interface keeps classifications tied to original conversational evidence and gives users direct control over how AI judgments are revised.Source: Fig. 4
  1. 01

    Product input and filtering

    Users specify a product or brand, initiate analysis, and use sentiment filters to explore collected content.

    Product input and filtering

    /research/messcleaner/fig-4-input.png

  2. 02

    Evaluation dimensions

    Candidate product-evaluation dimensions organize evidence; users select six and can replace or adjust them according to their priorities.

    Evaluation dimensions

    /research/messcleaner/fig-4-dimensions.png

Decision Synthesis

Reviewed evidence is carried forward into a product-level view rather than replaced by a new opaque summary.

Fig. 5

Authentic vs marketing / Consensus vs conflict / Radar / Risks / Personalized support

/research/messcleaner/fig-5-synthesis.png

Decision synthesis preserves the distinctions and uncertainty users have already inspected, allowing final judgments to remain grounded in traceable evidence.Source: Fig. 5

The interface compares evidence from authentic experience and suspicious marketing, highlights agreement and conflict, and summarizes common ground, key differences, and unresolved uncertainty. The overall view then integrates dimension-level findings into an evidence-based product view.

  1. Dimension-specific evidence

    Authentic experience and suspicious marketing remain visible within each selected product dimension.

  2. Consensus, conflict, and uncertainty

    Consistent findings, conflicting claims, and unresolved uncertainty are shown separately rather than smoothed away.

  3. Overall assessment and personalized support

    The overall assessment, radar chart, key findings and risks, and personalized decision support all rest on evidence the user has already reviewed.

Study Design

We compared MessCleaner with both unaided baseline practice and a generic AI-supported condition.

N
32
Design
Within-subjects
Conditions
Baseline / Basic AI / MessCleaner

Each participant completed one purchase decision-making task in all three conditions. Condition order was distributed across all six counterbalanced orders: ABC, ACB, BAC, BCA, CAB, CBA.

Fig. 6

Three parallel tracks, with counterbalancing shown as connecting paths

/research/messcleaner/fig-6-study.svg

Each condition included a purchase decision-making task and post-task questionnaire, followed by a final semi-structured interview; condition orders were counterbalanced across participants.Source: Fig. 6

Key Findings

Four questionnaire dimensions share the same seven-point scale, so conditions can be compared across the page.

D1 · eWOM Structuring and Evidence Organization

/research/messcleaner/fig-7-d1.svg

D2 · Credibility Judgment and Evidence Identification

/research/messcleaner/fig-7-d2.svg

Questionnaire results for eWOM Structuring and Evidence Organization and Credibility Judgment and Evidence Identification. Ratings use a seven-point Likert scale from Very low to Very high.Source: Fig. 7

D3 · AI Explanation, Transparency, and Control

/research/messcleaner/fig-8-d3.svg

D4 · eWOM Synthesis and Decision Quality

/research/messcleaner/fig-8-d4.svg

Questionnaire results for AI Explanation, Transparency, and Control and eWOM Synthesis and Decision Quality. Ratings use the same seven-point scale.Source: Fig. 8
  1. MessCleaner helped participants organize fragmented eWOM into structured and decision-relevant evidence and identify what information mattered.

    Evidence

    MessCleaner received higher ratings than both Basic AI and Baseline for Overall eWOM Understanding and Information Organization. For Valuable Information Identification, MessCleaner also outperformed both conditions, whereas Basic AI did not differ from Baseline.

    Interpretation

    Reducing information-processing effort and helping users determine what evidence matters are different forms of support; generic summarization did not consistently achieve the latter.

  2. MessCleaner supported users in distinguishing authentic experience, suspicious marketing, and uncertain information while inspecting and revising AI judgments.

    Evidence

    MessCleaner outperformed both Basic AI and Baseline on authentic/promotional distinction, promotional signal identification, and insufficient evidence recognition. It also received a higher Authenticity Judgment Confidence rating than both conditions while retaining an explicit uncertain-information category.

    Interpretation

    Confidence did not require hiding uncertainty; explicit uncertainty could signal where additional inspection was needed.

  3. Participants reported clearer understanding of AI assessments, more selective reliance, stronger perceived control, and stronger grounding of final decisions in reviewed evidence.

    Evidence

    Compared with Basic AI, MessCleaner received higher ratings for AI Judgment Understanding (6.13 vs 5.47), AI Reference-Dimension Visibility (6.56 vs 5.50), Trust Calibration, Avoiding Blind Acceptance, and Perceived Control. Decision Grounding: MessCleaner 6.19, Basic AI 5.19, Baseline 5.25. Decision Quality: MessCleaner 6.44, Basic AI 5.66, Baseline 4.88.

    Interpretation

    The system supported decisions by carrying forward evidence that users had already inspected and revised, rather than asking them to trust a detached AI conclusion.

Contribution

  1. Authenticity framework

    Reveal how human and AI authenticity judgments overlap and differ.

    Judgments from 40 consumers and six AI systems reveal shared, divergent, and complementary judgment dimensions.

  2. Collaborative judgment

    Turn AI authenticity assessments into inspectable and revisable objects.

    MessCleaner exposes AI reasoning, supporting evidence, and uncertainty so users can compare AI interpretations with their own and revise classifications.

  3. Decision synthesis

    Carry reviewed evidence into the final decision without replacing human judgment.

    The system synthesizes structured evidence while preserving its provenance, judgment differences, and uncertainty, keeping the user responsible for the final purchase decision.