[Part 2] How I Made 3 YouTube Shorts with ChatGPT and AI Agents
Building a One-Person AI Content Studio for $20/Month
I produced and published three AI-assisted YouTube Shorts. The first performance snapshot, measured through July 28, 2026, recorded 27 total views, 18 valid views, and no subscriber growth.
Those numbers did not prove success. They did reveal something more useful: the difficult part of AI Shorts production was not rendering a video. It was choosing a topic, verifying facts and rights, coordinating visuals, narration, captions, music, and timing, and then making a human quality decision.
![]() |
| The second production experiment in the AI one-person content studio: researching, producing, publishing, and measuring factual Shorts. |
The test used two factual stories and produced three language-specific uploads: a Korean and Chinese version of the North Pond hermit story, and a Korean version of the Balvano train disaster story.
What was actually published
The initial question was whether the same production structure could be repeated across topics and languages. Scripts, audio, captions, and video files could be generated. But the ability to produce a file did not prove audience interest or commercial potential.
The first challenge was choosing a topic
AI can generate long lists of news events, places, people, food, and historical incidents. A searchable or trending topic is not automatically a publishable topic.
- Would the topic create genuine curiosity for a specific viewer?
- Could the central facts be verified through reliable sources?
- Could I explain the usage rights for the images, footage, and music?
- Could the story be told visually in under a minute without repetitive scenes?
I also tested broad global trend research. The result was useful as an editorial candidate list, not as an objective popularity ranking. AI was good at expanding the candidate pool. The decision to publish remained an editorial responsibility.
A selected topic still did not produce a finished video
- the same person or place changed between scenes;
- backgrounds repeated while the central subject disappeared;
- pan-and-zoom motion felt artificial;
- narration, captions, and cuts drifted out of sync;
- changing the language changed the duration of the entire video.
Requests such as “make it more cinematic” were not reliable production instructions. I needed a visual storyboard, a purpose for each scene, and approved media before final assembly.
Free TTS and background music required human listening
Free and local text-to-speech tools were useful for comparing voices and estimating duration. Some voices still sounded mechanical, handled emotion poorly, or mispronounced names. Commercial-use rights also had to be checked separately from technical availability.
One Chinese North Pond version passed checks for audio presence, peaks, clipping, and ducking. When I listened to it, the track sounded more like white noise than useful background music.
The production roles
- Human: select the audience and topic, decide rights boundaries, judge perceived quality, and approve publication
- ChatGPT: organize research, compare topics, review facts, draft scripts, localize, and interpret KPI data
- Codex: assemble approved media, create production files, manage versions, and run technical checks
What the first KPI snapshot said
| Upload | Views | Valid views | Average percentage viewed | Subscribers |
|---|---|---|---|---|
| North Pond — Korean | 14 | 11 | 53.53% | 0 |
| Balvano — Korean | 8 | 4 | 61.92% | 0 |
| North Pond — Chinese | 5 | 3 | 47.26% | 0 |
| Total | 27 | 18 | Not combined | 0 |
The sample was too small to rank topics or languages. TikTok's observed view count was zero, but I did not combine it with YouTube because platform definitions and distribution systems are different.
Existing published evidence
These videos are linked in their original published languages. No separate English remake was created for this Blogger launch.
The realistic result: semi-automation
The most useful result was not full automation. It was a semi-automated workflow in which AI agents handled repeatable production and technical checks while a person made editorial, legal, and sensory decisions.
Twenty-seven views did not prove a business. They created the first operating data for deciding what could be automated, where human review was essential, and what to test next.
Part 3 applies the same approach to an original webtoon and motion comic distributed in multiple languages.
Previous: Part 1 — How I Built an AI-Powered Personal Asset Dashboard
Continue: Part 3 — How I Built a Multilingual AI Webtoon Workflow
This article is an English localization of a production record first published in Korean. Korean original / 한국어 원문.





Comments
Post a Comment