[Part 2] How I Made 3 YouTube Shorts with ChatGPT and AI Agents

Building a One-Person AI Content Studio for $20/Month

I produced and published three AI-assisted YouTube Shorts. The first performance snapshot, measured through July 28, 2026, recorded 27 total views, 18 valid views, and no subscriber growth.

Those numbers did not prove success. They did reveal something more useful: the difficult part of AI Shorts production was not rendering a video. It was choosing a topic, verifying facts and rights, coordinating visuals, narration, captions, music, and timing, and then making a human quality decision.

The second production experiment in the AI one-person content studio: researching, producing, publishing, and measuring factual Shorts.

The test used two factual stories and produced three language-specific uploads: a Korean and Chinese version of the North Pond hermit story, and a Korean version of the Balvano train disaster story.

What was actually published

Two factual stories became three language-specific uploads.

The initial question was whether the same production structure could be repeated across topics and languages. Scripts, audio, captions, and video files could be generated. But the ability to produce a file did not prove audience interest or commercial potential.

The first challenge was choosing a topic

AI can generate long lists of news events, places, people, food, and historical incidents. A searchable or trending topic is not automatically a publishable topic.

  1. Would the topic create genuine curiosity for a specific viewer?
  2. Could the central facts be verified through reliable sources?
  3. Could I explain the usage rights for the images, footage, and music?
  4. Could the story be told visually in under a minute without repetitive scenes?

I also tested broad global trend research. The result was useful as an editorial candidate list, not as an objective popularity ranking. AI was good at expanding the candidate pool. The decision to publish remained an editorial responsibility.

A selected topic still did not produce a finished video 

  • the same person or place changed between scenes;
  • backgrounds repeated while the central subject disappeared;
  • pan-and-zoom motion felt artificial;
  • narration, captions, and cuts drifted out of sync;
  • changing the language changed the duration of the entire video.

Requests such as “make it more cinematic” were not reliable production instructions. I needed a visual storyboard, a purpose for each scene, and approved media before final assembly.

Free TTS and background music required human listening

Free and local text-to-speech tools were useful for comparing voices and estimating duration. Some voices still sounded mechanical, handled emotion poorly, or mispronounced names. Commercial-use rights also had to be checked separately from technical availability.

One Chinese North Pond version passed checks for audio presence, peaks, clipping, and ducking. When I listened to it, the track sounded more like white noise than useful background music.

A technical pass did not mean the music felt right to a listener.

The production roles

  • Human: select the audience and topic, decide rights boundaries, judge perceived quality, and approve publication
  • ChatGPT: organize research, compare topics, review facts, draft scripts, localize, and interpret KPI data
  • Codex: assemble approved media, create production files, manage versions, and run technical checks
Human judgment, ChatGPT structure, and Codex execution were connected to real publishing data.

This is a practical version of AI-agent production and vibe coding: natural-language direction becomes files and checks, but responsibility for facts, rights, quality, and publication remains human.

What the first KPI snapshot said

UploadViewsValid viewsAverage percentage viewedSubscribers
North Pond — Korean141153.53%0
Balvano — Korean8461.92%0
North Pond — Chinese5347.26%0
Total2718Not combined0
A small sample is an input for the next hypothesis, not proof of success or failure.

The sample was too small to rank topics or languages. TikTok's observed view count was zero, but I did not combine it with YouTube because platform definitions and distribution systems are different.

Existing published evidence

These videos are linked in their original published languages. No separate English remake was created for this Blogger launch.

The realistic result: semi-automation

The most useful result was not full automation. It was a semi-automated workflow in which AI agents handled repeatable production and technical checks while a person made editorial, legal, and sensory decisions.

Twenty-seven views did not prove a business. They created the first operating data for deciding what could be automated, where human review was essential, and what to test next.

Part 3 applies the same approach to an original webtoon and motion comic distributed in multiple languages.

Previous: Part 1 — How I Built an AI-Powered Personal Asset Dashboard

Continue: Part 3 — How I Built a Multilingual AI Webtoon Workflow

This article is an English localization of a production record first published in Korean. Korean original / 한국어 원문.

Comments

Popular posts from this blog

[Prologue] How I Built a $20/Month AI Content Creation Workflow

[Part 1] How I Built an AI-Powered Personal Asset Dashboard