Designing
AI Quality Systems

Part 4: Authenticity Through Imperfection. Making AI Content Feel Human

December 18, 2025 · Kenneth Hung · 20 min read

Context: Generative AI solved the generation problem, not the quality problem, and in commerce video, quality is trust. Users detect AI content within seconds and swipe past it; when buyers don't trust the video, they don't buy. That trust gap gates GMV growth in APAC, US/EU, and LATAM: seller video supply, not demand, is the bottleneck, and AI-generated video only closes that gap if it converts like authentic content does.

Objective: Grow regional GMV by closing the trust gap between AI-generated and authentic seller video.

Goal: The AI Authenticity Framework, a reusable video generation system built from three fixed prompt layers (Style / Ceiling / Floor) plus variable narrative templates, that makes AI-generated seller video pass as human-created. The core architecture stays fixed across every region; language, background, and digital human casting adjust per market.

The Problem

Every video has to clear the same 3-second trust test buyers apply to real creator content. GMV growth depends on it.

Users don't trust AI-generated content, and trust is the foundation of TikTok Shop conversions.

AI-generated TikTok videos have five fatal flaws:

  1. Scripts sound like ads: grammatically perfect, logically linear, and start with the conclusion.

  2. Visuals are too clean: stable and centered framing, professional lighting, razor-sharp 4K.

  3. Performance is too perfect: confident throughout, flawlessly recited, with no pauses.

  4. Audio is too professional: silent recording environment, uniform volume levels, BGM synced precisely to the beat.

  5. Editing is too deliberate: fancy transitions, pinpoint subtitles, and an obvious ending.

The result:

  • Users identify it as an AI ad within 3 seconds → no trust → they swipe away.

Authenticity = Trust = GMV Growth

The Solution

The AI Authenticity Framework combines fixed architecture with variable templates:

  • Three-layer prompt architecture (fixed): Style / Ceiling / Floor – defining visual style, performance boundaries, and absolute no-go zones, applicable to all formats.

  • Multi-scenario narrative templates (variable): scene count and structure adjusted based on format and product category.

The result:

  • AI‑generated content feels human, not AI.

  • Content feels like it was "discovered," not "produced."

1: Why is AI-generated content instantly recognizable as AI?

Realism comes from imperfection.

AIGC content is instantly recognizable as AI not because the technology isn't good enough, but because every dimension is too perfect: scripts read like ad copy rather than thinking out loud; visuals look like professionally produced footage rather than casually recorded clips; performances feel like reciting lines rather than genuine, spontaneous reactions; audio sounds like a soundproof studio rather than an actual on‑site environment; and editing resembles polished commercials rather than rough cuts and raw assemblies.

Our goal: make content feel like it was discovered, not produced.

Problem 1: Scripts Sound Like Ads

Issue ❌ AI Style ✅ Human Style Solution Direction
Too linear, starts with the conclusion"This product is very innovative…""Okay wait... actually... hold on"Start with confusion, think out loud
Too much explanation, perfect grammarComplete, fluent sentences"I thought this would be dumb, but…"Interrupt sentences, leave things unfinished
Never contradicts itself"It's lightweight""Not super light... actually no, it is light"Add minor contradictions
Describes features"Breathable, lightweight material""Wore it for an hour and didn't feel like taking it off"Describe moments, not features
Uses marketing buzzwords"Optimized", "high‑quality", "innovative""Actually works", "doesn't feel cheap", "I was surprised"Replace marketing words with real human language
All pros, no cons"Perfectly solves all problems""The only thing I don't really like is here… but I'm being picky"Add an honest flaw to build trust
Doesn't acknowledge camera limitationsShows directly, assumes viewer can see clearly"Not sure if the camera can pick this up…"Add camera‑aware language
Opening sounds like a CTA"Today I'm introducing to everyone…""I wasn't planning to post this…"Use an unexpected hook, not an intro
Issue ❌ AI Style ✅ Human Style Solution Direction
Camera too stableGimbal‑smooth, fluid motionSlight handheld shake, unstableShoot handheld, keep natural shake
Composition perfectly centeredSubject dead center, perfectly framedSubject off‑center, head occasionally out of frameIntentionally imperfect framing
Professional lightingStudio‑grade lighting, no shadowsWindow sidelight, overhead warm lamp, phone + ring lightUse natural indoor light, accept uneven exposure
Image too sharp4K razor‑sharp, no noiseMobile compression feel, shadow noise, soft focusReduce sharpness, add compression feel
Focus perfectAlways locked in sharp focusSlight focus hunting, auto‑exposure adjustmentsAllow focus and exposure to auto‑adjust
No compression artifactsRaw quality, no encoding lossTikTok‑style compression, rolling shutterDeliberately add compression and slight rolling shutter effect
No platform interaction signalsPure recording, no interferenceAdjusting phone, glancing at screen, reacting to off‑screen events while recordingAdd actions like adjusting the device, looking at screen
Overall feel"This was professionally shot""I didn't plan to post this, just casually recorded it"Goal is "casual recording" rather than "professional production"
Issue ❌ AI Style ✅ Human Style Solution Direction
EnergyMonotone throughoutVaried, emphasizes only at key pointsAdd vocal variety, emphasize at key moments
Feels rehearsedSounds like reading linesSounds like thinking out loudSimulate thought process, not recitation
No pausesComplete, uninterrupted sentences"Um", "like", sentence restartsAdd filler words, pauses, restarts
Reacts too quicklyImmediate, precise reactionsSlightly delayed, occasionally awkwardAllow reaction delays and awkward moments
Eyes fixed on cameraConstant eye contact with lensLooks at screen, side glances, adjusts phoneAdd screen glances and device adjustments
Expressions stiff or over‑exaggeratedPerformative, deliberate expressionsNatural blinks, micro‑expressions, facial asymmetryNatural micro‑expressions, avoid performance feel
Movements too smoothEvery move purposefulClumsy movement, unsure where to put handsAllow aimless small movements and awkwardness
Issue ❌ AI Style ✅ Human Style Solution Direction
Ambient soundStudio‑silent, completely isolatedHas ambient noise, slight echo, room feelRetain ambient sound and room reverb
Volume uniformityStable, professionally processedUneven volume, occasional sudden changesDon't normalize volume
MicrophoneProfessional mic qualityPhone mic quality, slight background noiseUse phone microphone quality
Background musicLicensed music synced to beat, music changes at endNo background music, or stays consistent throughoutNo music, or keep it consistent
Overall textureClean, sounds like a podcastSounds like it was casually recorded in a roomGoal is "casual recording" rather than "professional production"
Issue ❌ AI Style ✅ Human Style Solution Direction
Transitions too smoothFades, fancy transitionsHard cuts, no transitionsUse only hard cuts, no transitions
Captions too preciseCaptions appear perfectly in syncCaptions slightly delayed, like they were added laterDelay captions by 0.2–0.5 seconds
Rhythm feels cinematicRhythmic, beat‑synced, structured arcCasual editing, jump cuts mid‑sentenceJump cut mid‑sentence, don't worry about rhythm
Has a sense of closureClear ending, music fades outAbrupt ending, like recording stopped mid‑takeDon't wrap up — just stop

2: Three-Layer Prompt Architecture

Before generating any scene, set a global prompt that defines the video's tone and constraints. These three layers are stacked together to ensure that AI-generated content never pushes beyond the boundaries of realism.

Floor
Defines absolute no-go zones
Believable everyday behavior Avoid marketing language Avoid perfect grammar Avoid influencer or commercial tone
Ceiling
Defines upper performance limit
Unrehearsed, thinking-out-loud delivery Pauses, restarts, filler words Natural blinking and facial micro-expressions Slightly unsure at first, then gradually confident
Style
Defines visual style
Casual TikTok-native phone video Handheld smartphone camera, slight natural shake Imperfect framing, subject slightly off-center Indoor natural lighting, uneven exposure No studio lighting, no cinematic look Mobile video compression, soft focus Slight digital noise and motion blur Looks like a normal person filming at home Natural skin texture, minimal makeup

*Here, Floor / Ceiling / Style apply to behavioral realism (a video's performance, delivery, and visual texture) rather than content category, which is how Part 1 originally defined them. Same three-layer structure, applied one level down: to how a single video behaves, not what kind of video it is.

3: Multi‑Scenario Narrative Template

Core principle: multiple short scenes, rather than one long continuous shot. Scene count and structure vary by narrative format.

An effective TikTok seller video isn't about "introducing a product" — it's about telling a story of discovery.

5-Scene Narrative · Problem → Solution

Click any scene card to explore the visual prompt, script, humanization rules, and quality flags for that step of the arc.

Relatability
Credibility
Discovery
Trust
Conversion
0s 7s 14s 24s 35s 45s
Style
Floor
Style Ceiling
Ceiling
Floor
1
Problem / Hook
0–7s · Relatability
Style

Primary stake. All three layers still execute below.

Visual Prompt
Defines the Style layer for this scene

Selfie camera turns on mid-motion. Creator adjusts phone angle while already talking. Head partially cropped for a moment. Phone slips slightly or tilts as she moves it, showing subtle frustration without reacting dramatically. Natural, slightly distracted body language.

Script
Written per the Humanization Rules below

"Okay this might just be me, but does anyone else hate when your phone just… won't stand up anywhere?"

Humanization Rules
Governs the Script above, raises output toward the Ceiling
Start with uncertainty Camera-aware language
Regenerate If
Signals a Floor violation, redo this take
Feels too clean First 2 seconds feel too deliberate Composition too perfect / centered
2
Failed Solutions
7–14s · Credibility
Floor

Primary stake. All three layers still execute below.

Visual Prompt
Defines the Style layer for this scene

Hard jump cut feeling. Same person, same room, same outfit. Creator looks at the phone screen, not directly at the lens. She casually flips her phone around and tries to rest it against a couple of bulky phone stands or objects. The phone looks awkward and slightly unstable. She doesn't try to make it work perfectly. Natural, slightly annoyed body language. Small pause while adjusting the phone, then giving up.

Script
Written per the Humanization Rules below

"I've tried propping it on stuff, and those big stands are annoying, and I didn't wanna stick something bulky on my phone."

Humanization Rules
Governs the Script above, raises output toward the Ceiling
Interrupt sentences Limit vocabulary
Regenerate If
Signals a Floor violation, redo this take
Too energetic Sounds like listing features Eye contact with camera
3
Solution Discovery
14–24s · Discovery
Style Ceiling

Primary stake. All three layers still execute below.

Visual Prompt
Defines the Style layer for this scene

Soft jump cut. She casually flips her phone over, revealing a compact magnetic phone stand already attached to the back. She doesn't explain it immediately. Small pause as she adjusts the phone slightly.

Script
Written per the Humanization Rules below

"Someone sent me this little magnetic thing, and I honestly thought it was kind of pointless… but I've been using it for like two days."

Humanization Rules
Governs the Script above, raises output toward the Ceiling
Minor contradiction Start with uncertainty
Regenerate If
Signals a Floor violation, redo this take
Product is lit or floating Product perfectly centered Sounds excited about the product
4
Proof / Trust
24–35s · Trust
Ceiling

Primary stake. All three layers still execute below.

Visual Prompt
Defines the Style layer for this scene

Hard jump cut. She sets the phone down on the desk using the stand while continuing to talk, not stopping to show it. She briefly looks at the screen to check framing, then continues naturally.

Script
Written per the Humanization Rules below

"This is what sold me, when I'm on FaceTime or watching something while eating, it just works. I mean, it's not the prettiest thing, but it folds flat so I forget it's there."

Humanization Rules
Governs the Script above, raises output toward the Ceiling
Moments > Features One honest flaw Minor contradiction
Regenerate If
Signals a Floor violation, redo this take
All praise, no flaw mentioned Generic benefit description Uses marketing language
5
Decision + CTA
35–45s · Conversion
Floor

Primary stake. All three layers still execute below.

Visual Prompt
Defines the Style layer for this scene

Same framing as previous scenes. She casually picks the phone back up. The compact magnetic stand folds flat naturally as she grabs it. No pause, no emphasis, no change in posture or energy. After speaking, she holds the phone naturally for 1–2 seconds, subtly glancing at the screen or moving slightly. Camera feels handheld and imperfect. Looks like a real moment caught on camera, not a demonstration.

Script
Written per the Humanization Rules below

"People kept asking where I got it, so I just dropped the link here. If you're dealing with the same thing, it's honestly pretty useful."

Humanization Rules
Governs the Script above, raises output toward the Ceiling
Limit vocabulary Interrupt sentences
Regenerate If
Signals a Floor violation, redo this take
Says "buy now" Contains urgency / scarcity Discount language Music or graphic change

The template above is the anatomy of the Problem → Solution narrative format: 5 scenes, each broken into a visual prompt, script, humanization rules, and regenerate-if flags.

Other narrative formats range from 3 to 6 scenes, but the core principle holds: multiple short scenes, not one continuous take. The three-layer prompt architecture applies to every format, though which layer takes primary stake shifts scene to scene. See the Template Library in Section 4 below for other examples.

Production disclosure

This is an exploratory experiment and I welcome your criticism and suggestions. This video was generated using Google Flow / Veo 3.1 / Nano Banana Pro to test whether the three-layer prompt architecture combined with a five-scene narrative structure can produce realistic AIGC video content.

The framework proves workable, but it also exposes known limitations: over‑softened skin textures, overly smooth hands lacking natural surface detail, and inaccurate physical interactions between products and the human body, among other issues. For more details, see the 'Technical Insights and Solutions' section below.

4: Can this framework scale?

The AI Authenticity Framework is built on a "fixed architecture + variable templates" model: the three-layer prompt system, chain-of-reference, and realism principles remain constant, while narrative format, scene count, product category, and target region can all be adjusted.

In theory, this scales, but turning it into reality needs real dialogue between product and engineering. I'm sharing this as a starting point, not a finished answer. If you are working in the same space, I'd love to connect and discuss.

The Framework

Architecture Layering

LayerComponentDescription
Fixed Three-Layer Prompts
🎨 Style 📈 Ceiling 🚫 Floor
Global constraint, prepended to every scene
Chain-of-Reference Hero Image → each scene inherits from the last
Realism + Assembly Imperfection = Realism, Hard Cuts, Mobile Texture
Variable Category / Region Beauty, 3C, Apparel… × US, SEA, CN…
Narrative Format Problem/Solution / Unboxing / Review / GRWM
9 examples below ↓
Narrative Format, instantiated as 9 templates

Template Library Examples

💡 Problem → Solution 5 beats · 5-7 scenes
Pain Point → Failure → Discovery → Proof → CTA
🔄 Before / After 3 beats · 5-6 scenes
Before (1-2) → Application (1-2) → After (1-2)
📦 Unboxing / First Impressions 4 beats · 5-6 scenes
Receiving (1) → Unboxing (1-2) → Reaction (1-2) → Conclusion (1)
🔍 Honest Review / Test 4 beats · 5-6 scenes
Skepticism (1) → Testing (2) → Pros/Cons (1-2) → Conclusion (1)
💄 GRWM 3 beats · 4-6 scenes
Prep (1-2) → Application (2-3) → Ready (1)
📚 Tutorial 4 beats · 5-7 scenes
Hook (1) → Step-by-Step Demo (2-3) → Result (1-2) → CTA (1)
🎧 ASMR 3 beats · 4-5 scenes
Sound Intro (1) → Sensory Close-ups (2-3) → Satisfying Payoff (1)
🎤 Vox Pop 4 beats · 5-6 scenes
Question Hook (1) → Testimonials (2-3) → Consensus (1) → CTA (1)
🎭 Skit 4 beats · 5-6 scenes
Setup (1) → Escalation (1-2) → Twist (1-2) → Resolution/CTA (1)
Fixed Architecture + Variable Templates = Scalable Framework

5. AIGC Video Generation: Technical Insights & Solutions

At the generation level: every fix below exists to close one specific gap between AI output and what a real seller would actually produce; the video-level trust test above is only as strong as the generation pipeline underneath it.

Digital Human Authenticity

Solution: Make people look like real humans, not AI‑generated fakes, and avoid that "uncanny AI human" feel.

🏷️ Issue Type ⚠️ Observed Issue 🔍 Root Cause ✅ Solution
Character DriftEach regeneration looks like a different personUnder‑specified text prompts, no visual anchor → high varianceEstablish a Hero Image as single source of truth; generate 3‑6 angle variants
Over‑smoothingSkin appears softened, filtered, unnatural textureTraining data biased toward beauty‑retouched imagesExplicitly request: visible pores, uneven skin tone, minor blemishes, camera noise
Gender BiasHands default to masculine features even when subject is femaleModel optimized for a "generic human"Specify finger thickness, nail length, skin texture + negative constraints
Anatomical ErrorsWrong number of fingers, unnatural joint anglesModel's limited understanding of human anatomyAvoid close‑up hand shots; substitute with isolated product shots
Unnatural ExpressionsStiff, over‑perfect, or overly exaggerated expressionsModel tends to generate "standard" expressionsRequest natural micro‑expressions, facial asymmetry, subtle blinking
Over‑rehearsed PerformanceConfident and fluent throughout, sounds like reading a scriptModel leans toward "perfect" deliveryStart with uncertainty, include pauses and restarts, allow awkward moments
Excessive PolishLooks professionally shot, not UGCModel defaults to high‑quality, stable visualsRequire handheld shake, imperfect framing, natural light, mobile compression feel
Over‑ExplainingActions are narrated, making it feel like an adUsers tolerate visual noise but reject overt sales intentKeep dialogue casual and incomplete; let visuals "accidentally" demonstrate
Reference MaterialUsing brand accounts or top‑creator content as referencesThat content is "produced," not "discovered"Reference real‑user content in the 5–50K view range
Flaws as FeaturesCompression artifacts, slight noise, rolling shutterThese "defects" are a byproduct of real‑world recordingEmbrace certain artifacts as part of UGC texture; realism comes from imperfection
🏷️ Issue Type ⚠️ Observed Issue 🔍 Root Cause ✅ Solution
Environment DriftEach regeneration has a different room/backgroundModel does not "remember" previous scenesLock environment into the Hero Image; all subsequent scenes inherit the same environment
Inconsistent LightingLighting direction, color temp, and intensity shift between scenesModel infers lighting independently each timeKeep lighting direction consistent in the reference set; each scene inherits the previous scene's lighting
Inconsistent Camera LogicCamera angle and height jump between scenesModel infers camera position independently each timeKeep camera position consistent in the reference set; chain‑of‑reference inherits camera logic
Scene DiscontinuityConsecutive scenes feel like different videos stitched togetherEach scene is generated independently, with no inheritance relationshipChain‑of‑Reference: each scene's input includes the selected image from the previous scene
Drift AccumulationLater scenes drift further from the opening sceneSmall errors from each scene accumulate over timeEach scene inputs both the original reference set + the previous scene's image
🏷️ Issue Type ⚠️ Observed Issue 🔍 Root Cause ✅ Solution
Product HallucinationProduct appearance, details, and color differ from realityModel "invents" the product rather than referencing real imagesPeople are generated, products are referenced — always use clean product reference images
Product DriftProduct looks different across scenesEach scene independently generates the product's appearancePrepare white‑background multi‑angle product shots as ground truth
Forced Product PlacementProduct appears abruptly, feeling like a hard sellNo deliberate pacing for product introductionGradual reveal: hint / off‑screen → partial interaction → clear display
Physical Interaction ErrorsPhone + stand interactions fail, fold points are guessed incorrectlyModel does not simulate physics, only interpolates visuallySplit into pre‑action → mid‑action → post‑action states; supplement with product‑only shots

Thank You!

This closes the four-part series: framework → scale → cross-regional practice → authenticity in practice.
If you're building something similar, I'd love to hear from you.