Designing
AI Quality Systems
Part 4: Authenticity Through Imperfection. Making AI Content Feel Human
December 18, 2025 · Kenneth Hung · 20 min read
Context: Generative AI solved the generation problem, not the quality problem, and in commerce video, quality is trust. Users detect AI content within seconds and swipe past it; when buyers don't trust the video, they don't buy. That trust gap gates GMV growth in APAC, US/EU, and LATAM: seller video supply, not demand, is the bottleneck, and AI-generated video only closes that gap if it converts like authentic content does.
Objective: Grow regional GMV by closing the trust gap between AI-generated and authentic seller video.
Goal: The AI Authenticity Framework, a reusable video generation system built from three fixed prompt layers (Style / Ceiling / Floor) plus variable narrative templates, that makes AI-generated seller video pass as human-created. The core architecture stays fixed across every region; language, background, and digital human casting adjust per market.
-
Part 1 introduced the Floor / Ceiling / Style framework.
Part 2 covered how that framework scales across verticals, markets, and execution, and introduced AI Behavior Design as a practice.
Part 3 puts the framework into practice for the cross-regional adaptation use case: one product, three markets, one quality system.
Part 4 (this piece) applies the same framework to a different use case, the AI Authenticity Framework, for making AI-generated content feel like authentic human creation, in service of regional GMV growth.
-
The Problem
Every video has to clear the same 3-second trust test buyers apply to real creator content. GMV growth depends on it.
Users don't trust AI-generated content, and trust is the foundation of TikTok Shop conversions.
AI-generated TikTok videos have five fatal flaws:
Scripts sound like ads: grammatically perfect, logically linear, and start with the conclusion.
Visuals are too clean: stable and centered framing, professional lighting, razor-sharp 4K.
Performance is too perfect: confident throughout, flawlessly recited, with no pauses.
Audio is too professional: silent recording environment, uniform volume levels, BGM synced precisely to the beat.
Editing is too deliberate: fancy transitions, pinpoint subtitles, and an obvious ending.
The result:
Users identify it as an AI ad within 3 seconds → no trust → they swipe away.
Authenticity = Trust = GMV Growth
The Solution
The AI Authenticity Framework combines fixed architecture with variable templates:
Three-layer prompt architecture (fixed): Style / Ceiling / Floor – defining visual style, performance boundaries, and absolute no-go zones, applicable to all formats.
Multi-scenario narrative templates (variable): scene count and structure adjusted based on format and product category.
The result:
AI‑generated content feels human, not AI.
Content feels like it was "discovered," not "produced."
1: Why is AI-generated content instantly recognizable as AI?
Realism comes from imperfection.
AIGC content is instantly recognizable as AI not because the technology isn't good enough, but because every dimension is too perfect: scripts read like ad copy rather than thinking out loud; visuals look like professionally produced footage rather than casually recorded clips; performances feel like reciting lines rather than genuine, spontaneous reactions; audio sounds like a soundproof studio rather than an actual on‑site environment; and editing resembles polished commercials rather than rough cuts and raw assemblies.
Our goal: make content feel like it was discovered, not produced.
Problem 1: Scripts Sound Like Ads
| Issue | ❌ AI Style | ✅ Human Style | Solution Direction |
|---|---|---|---|
| Too linear, starts with the conclusion | "This product is very innovative…" | "Okay wait... actually... hold on" | Start with confusion, think out loud |
| Too much explanation, perfect grammar | Complete, fluent sentences | "I thought this would be dumb, but…" | Interrupt sentences, leave things unfinished |
| Never contradicts itself | "It's lightweight" | "Not super light... actually no, it is light" | Add minor contradictions |
| Describes features | "Breathable, lightweight material" | "Wore it for an hour and didn't feel like taking it off" | Describe moments, not features |
| Uses marketing buzzwords | "Optimized", "high‑quality", "innovative" | "Actually works", "doesn't feel cheap", "I was surprised" | Replace marketing words with real human language |
| All pros, no cons | "Perfectly solves all problems" | "The only thing I don't really like is here… but I'm being picky" | Add an honest flaw to build trust |
| Doesn't acknowledge camera limitations | Shows directly, assumes viewer can see clearly | "Not sure if the camera can pick this up…" | Add camera‑aware language |
| Opening sounds like a CTA | "Today I'm introducing to everyone…" | "I wasn't planning to post this…" | Use an unexpected hook, not an intro |
| Issue | ❌ AI Style | ✅ Human Style | Solution Direction |
|---|---|---|---|
| Camera too stable | Gimbal‑smooth, fluid motion | Slight handheld shake, unstable | Shoot handheld, keep natural shake |
| Composition perfectly centered | Subject dead center, perfectly framed | Subject off‑center, head occasionally out of frame | Intentionally imperfect framing |
| Professional lighting | Studio‑grade lighting, no shadows | Window sidelight, overhead warm lamp, phone + ring light | Use natural indoor light, accept uneven exposure |
| Image too sharp | 4K razor‑sharp, no noise | Mobile compression feel, shadow noise, soft focus | Reduce sharpness, add compression feel |
| Focus perfect | Always locked in sharp focus | Slight focus hunting, auto‑exposure adjustments | Allow focus and exposure to auto‑adjust |
| No compression artifacts | Raw quality, no encoding loss | TikTok‑style compression, rolling shutter | Deliberately add compression and slight rolling shutter effect |
| No platform interaction signals | Pure recording, no interference | Adjusting phone, glancing at screen, reacting to off‑screen events while recording | Add actions like adjusting the device, looking at screen |
| Overall feel | "This was professionally shot" | "I didn't plan to post this, just casually recorded it" | Goal is "casual recording" rather than "professional production" |
| Issue | ❌ AI Style | ✅ Human Style | Solution Direction |
|---|---|---|---|
| Energy | Monotone throughout | Varied, emphasizes only at key points | Add vocal variety, emphasize at key moments |
| Feels rehearsed | Sounds like reading lines | Sounds like thinking out loud | Simulate thought process, not recitation |
| No pauses | Complete, uninterrupted sentences | "Um", "like", sentence restarts | Add filler words, pauses, restarts |
| Reacts too quickly | Immediate, precise reactions | Slightly delayed, occasionally awkward | Allow reaction delays and awkward moments |
| Eyes fixed on camera | Constant eye contact with lens | Looks at screen, side glances, adjusts phone | Add screen glances and device adjustments |
| Expressions stiff or over‑exaggerated | Performative, deliberate expressions | Natural blinks, micro‑expressions, facial asymmetry | Natural micro‑expressions, avoid performance feel |
| Movements too smooth | Every move purposeful | Clumsy movement, unsure where to put hands | Allow aimless small movements and awkwardness |
| Issue | ❌ AI Style | ✅ Human Style | Solution Direction |
|---|---|---|---|
| Ambient sound | Studio‑silent, completely isolated | Has ambient noise, slight echo, room feel | Retain ambient sound and room reverb |
| Volume uniformity | Stable, professionally processed | Uneven volume, occasional sudden changes | Don't normalize volume |
| Microphone | Professional mic quality | Phone mic quality, slight background noise | Use phone microphone quality |
| Background music | Licensed music synced to beat, music changes at end | No background music, or stays consistent throughout | No music, or keep it consistent |
| Overall texture | Clean, sounds like a podcast | Sounds like it was casually recorded in a room | Goal is "casual recording" rather than "professional production" |
| Issue | ❌ AI Style | ✅ Human Style | Solution Direction |
|---|---|---|---|
| Transitions too smooth | Fades, fancy transitions | Hard cuts, no transitions | Use only hard cuts, no transitions |
| Captions too precise | Captions appear perfectly in sync | Captions slightly delayed, like they were added later | Delay captions by 0.2–0.5 seconds |
| Rhythm feels cinematic | Rhythmic, beat‑synced, structured arc | Casual editing, jump cuts mid‑sentence | Jump cut mid‑sentence, don't worry about rhythm |
| Has a sense of closure | Clear ending, music fades out | Abrupt ending, like recording stopped mid‑take | Don't wrap up — just stop |
2: Three-Layer Prompt Architecture
Before generating any scene, set a global prompt that defines the video's tone and constraints. These three layers are stacked together to ensure that AI-generated content never pushes beyond the boundaries of realism.
*Here, Floor / Ceiling / Style apply to behavioral realism (a video's performance, delivery, and visual texture) rather than content category, which is how Part 1 originally defined them. Same three-layer structure, applied one level down: to how a single video behaves, not what kind of video it is.
3: Multi‑Scenario Narrative Template
Core principle: multiple short scenes, rather than one long continuous shot. Scene count and structure vary by narrative format.
An effective TikTok seller video isn't about "introducing a product" — it's about telling a story of discovery.
5-Scene Narrative · Problem → Solution
Click any scene card to explore the visual prompt, script, humanization rules, and quality flags for that step of the arc.
Selfie camera turns on mid-motion. Creator adjusts phone angle while already talking. Head partially cropped for a moment. Phone slips slightly or tilts as she moves it, showing subtle frustration without reacting dramatically. Natural, slightly distracted body language.
"Okay this might just be me, but does anyone else hate when your phone just… won't stand up anywhere?"
Hard jump cut feeling. Same person, same room, same outfit. Creator looks at the phone screen, not directly at the lens. She casually flips her phone around and tries to rest it against a couple of bulky phone stands or objects. The phone looks awkward and slightly unstable. She doesn't try to make it work perfectly. Natural, slightly annoyed body language. Small pause while adjusting the phone, then giving up.
"I've tried propping it on stuff, and those big stands are annoying, and I didn't wanna stick something bulky on my phone."
Soft jump cut. She casually flips her phone over, revealing a compact magnetic phone stand already attached to the back. She doesn't explain it immediately. Small pause as she adjusts the phone slightly.
"Someone sent me this little magnetic thing, and I honestly thought it was kind of pointless… but I've been using it for like two days."
Hard jump cut. She sets the phone down on the desk using the stand while continuing to talk, not stopping to show it. She briefly looks at the screen to check framing, then continues naturally.
"This is what sold me, when I'm on FaceTime or watching something while eating, it just works. I mean, it's not the prettiest thing, but it folds flat so I forget it's there."
Same framing as previous scenes. She casually picks the phone back up. The compact magnetic stand folds flat naturally as she grabs it. No pause, no emphasis, no change in posture or energy. After speaking, she holds the phone naturally for 1–2 seconds, subtly glancing at the screen or moving slightly. Camera feels handheld and imperfect. Looks like a real moment caught on camera, not a demonstration.
"People kept asking where I got it, so I just dropped the link here. If you're dealing with the same thing, it's honestly pretty useful."
The template above is the anatomy of the Problem → Solution narrative format: 5 scenes, each broken into a visual prompt, script, humanization rules, and regenerate-if flags.
Other narrative formats range from 3 to 6 scenes, but the core principle holds: multiple short scenes, not one continuous take. The three-layer prompt architecture applies to every format, though which layer takes primary stake shifts scene to scene. See the Template Library in Section 4 below for other examples.
Production disclosure
This is an exploratory experiment and I welcome your criticism and suggestions. This video was generated using Google Flow / Veo 3.1 / Nano Banana Pro to test whether the three-layer prompt architecture combined with a five-scene narrative structure can produce realistic AIGC video content.
The framework proves workable, but it also exposes known limitations: over‑softened skin textures, overly smooth hands lacking natural surface detail, and inaccurate physical interactions between products and the human body, among other issues. For more details, see the 'Technical Insights and Solutions' section below.
4: Can this framework scale?
The AI Authenticity Framework is built on a "fixed architecture + variable templates" model: the three-layer prompt system, chain-of-reference, and realism principles remain constant, while narrative format, scene count, product category, and target region can all be adjusted.
In theory, this scales, but turning it into reality needs real dialogue between product and engineering. I'm sharing this as a starting point, not a finished answer. If you are working in the same space, I'd love to connect and discuss.
Architecture Layering
| Layer | Component | Description |
|---|---|---|
| Fixed | Three-Layer Prompts |
🎨 Style
📈 Ceiling
🚫 Floor
Global constraint, prepended to every scene
|
| Chain-of-Reference | Hero Image → each scene inherits from the last | |
| Realism + Assembly | Imperfection = Realism, Hard Cuts, Mobile Texture | |
| Variable | Category / Region | Beauty, 3C, Apparel… × US, SEA, CN… |
| Narrative Format | Problem/Solution / Unboxing / Review / GRWM 9 examples below ↓ |
Template Library Examples
5. AIGC Video Generation: Technical Insights & Solutions
At the generation level: every fix below exists to close one specific gap between AI output and what a real seller would actually produce; the video-level trust test above is only as strong as the generation pipeline underneath it.
Digital Human Authenticity
Solution: Make people look like real humans, not AI‑generated fakes, and avoid that "uncanny AI human" feel.
| 🏷️ Issue Type | ⚠️ Observed Issue | 🔍 Root Cause | ✅ Solution |
|---|---|---|---|
| Character Drift | Each regeneration looks like a different person | Under‑specified text prompts, no visual anchor → high variance | Establish a Hero Image as single source of truth; generate 3‑6 angle variants |
| Over‑smoothing | Skin appears softened, filtered, unnatural texture | Training data biased toward beauty‑retouched images | Explicitly request: visible pores, uneven skin tone, minor blemishes, camera noise |
| Gender Bias | Hands default to masculine features even when subject is female | Model optimized for a "generic human" | Specify finger thickness, nail length, skin texture + negative constraints |
| Anatomical Errors | Wrong number of fingers, unnatural joint angles | Model's limited understanding of human anatomy | Avoid close‑up hand shots; substitute with isolated product shots |
| Unnatural Expressions | Stiff, over‑perfect, or overly exaggerated expressions | Model tends to generate "standard" expressions | Request natural micro‑expressions, facial asymmetry, subtle blinking |
| Over‑rehearsed Performance | Confident and fluent throughout, sounds like reading a script | Model leans toward "perfect" delivery | Start with uncertainty, include pauses and restarts, allow awkward moments |
| Excessive Polish | Looks professionally shot, not UGC | Model defaults to high‑quality, stable visuals | Require handheld shake, imperfect framing, natural light, mobile compression feel |
| Over‑Explaining | Actions are narrated, making it feel like an ad | Users tolerate visual noise but reject overt sales intent | Keep dialogue casual and incomplete; let visuals "accidentally" demonstrate |
| Reference Material | Using brand accounts or top‑creator content as references | That content is "produced," not "discovered" | Reference real‑user content in the 5–50K view range |
| Flaws as Features | Compression artifacts, slight noise, rolling shutter | These "defects" are a byproduct of real‑world recording | Embrace certain artifacts as part of UGC texture; realism comes from imperfection |
| 🏷️ Issue Type | ⚠️ Observed Issue | 🔍 Root Cause | ✅ Solution |
|---|---|---|---|
| Environment Drift | Each regeneration has a different room/background | Model does not "remember" previous scenes | Lock environment into the Hero Image; all subsequent scenes inherit the same environment |
| Inconsistent Lighting | Lighting direction, color temp, and intensity shift between scenes | Model infers lighting independently each time | Keep lighting direction consistent in the reference set; each scene inherits the previous scene's lighting |
| Inconsistent Camera Logic | Camera angle and height jump between scenes | Model infers camera position independently each time | Keep camera position consistent in the reference set; chain‑of‑reference inherits camera logic |
| Scene Discontinuity | Consecutive scenes feel like different videos stitched together | Each scene is generated independently, with no inheritance relationship | Chain‑of‑Reference: each scene's input includes the selected image from the previous scene |
| Drift Accumulation | Later scenes drift further from the opening scene | Small errors from each scene accumulate over time | Each scene inputs both the original reference set + the previous scene's image |
| 🏷️ Issue Type | ⚠️ Observed Issue | 🔍 Root Cause | ✅ Solution |
|---|---|---|---|
| Product Hallucination | Product appearance, details, and color differ from reality | Model "invents" the product rather than referencing real images | People are generated, products are referenced — always use clean product reference images |
| Product Drift | Product looks different across scenes | Each scene independently generates the product's appearance | Prepare white‑background multi‑angle product shots as ground truth |
| Forced Product Placement | Product appears abruptly, feeling like a hard sell | No deliberate pacing for product introduction | Gradual reveal: hint / off‑screen → partial interaction → clear display |
| Physical Interaction Errors | Phone + stand interactions fail, fold points are guessed incorrectly | Model does not simulate physics, only interpolates visually | Split into pre‑action → mid‑action → post‑action states; supplement with product‑only shots |
Thank You!
This closes the four-part series: framework → scale → cross-regional practice → authenticity in practice.
If you're building something similar, I'd love to hear from you.