# Plan and read package tests

Use this for a title, thumbnail, or paired test. An A/B test shows different options to viewers and compares the results.

## Define the question

Ask what would change after the test. Choose one question that matters enough to spend traffic on. Record the current pair as the control, meaning the version used as a baseline.

For a new promise, test whole pairs. A winning pair tells you which package worked. It does not tell you whether the title or image caused the gain. For a narrower lesson, hold one part fixed. Keep each option in the show's visual system.

Do not prescribe the same rounds for every episode. Start with the weak point. If the image is settled, test titles first. If all titles work with the same image, a title test may be enough.

## Check the tool

Verify current YouTube Help and the user's account before giving setup steps. The rules below were checked on September 28, 2026.

YouTube Studio, the site's creator dashboard, supports up to three titles, thumbnails, or paired options per test. It requires desktop access and advanced features. It excludes Shorts, private videos, videos made for kids, age-restricted videos, scheduled live streams, and active Premieres. Replays can be eligible. A test can take up to two weeks. Changing a live title or thumbnail stops it. Read its watch-time result and reported confidence. A small lead in the numbers is not enough.

The tool can report Winner, Performed the same, or Inconclusive. Without a clear winner, the first uploaded option becomes the default. Check the exact labels in the report.

If testing is unavailable, give a ready-to-run plan and pick an editorial default. Do not pretend manual swaps are a controlled test.

For another tool, check whether variants run at the same time or in sequence. Check its winning metric. A tool that rotates options across hours can mix an option's effect with changes in viewers or traffic.

## Make the plan

Return one table:

| Round | Question | Options and asset IDs | Held fixed | Tool and measure | End rule |
| --- | --- | --- | --- | --- | --- |
| 1 | Which promise earns better viewing? | Current pair A, candidate B, candidate C | Episode cut and style rules | Native paired test; reported watch-time result | Wait for completion; act sooner only for a false claim or broken asset |

Add the planned start and review date. Note any guest promotion, paid traffic, or major news event that could change the audience. Avoid overlapping experiments on the same package.

Use staged rounds only when the first result leaves a useful question. For example, choose between two base shots, then test text on the selected shot. Keep the best supported option as the next baseline. Do not send 20 design drafts into 20 live rounds by default.

## Read the evidence

Copy the report's result, dates, option IDs, and numbers. Separate the test's numbers from whole-episode numbers.

Define terms when used:

- Impressions: counted times a thumbnail appeared on YouTube.
- Click-through rate, or CTR: the share of those impressions that led to a view.
- Watch time: time viewers spent watching.
- Watch-time share: the portion of the test's watch time credited to an option.
- Retention: how much of the audience stays through each part of the video.

Use the report's result to choose the package. Use the episode's traffic and viewing data to frame a follow-up question. Do not assign the episode's CTR or retention to one option unless the tool provides that split.

A hypothetical report shows A at 55% of watch time and B at 45%, with an Inconclusive label. Report no proven winner. Do not override that label with the percentages. If the tool reports A as Winner, choose A for this test; the report still does not show why B lost.

Avoid turning one result into a rule for the whole show. Later viewers may differ from early fans. Review a pattern across several episodes before adding it to the show profile as a lesson.

## Decide when to stop

- Keep a clear winner if it meets the show's rules and the episode keeps its promise.
- When options perform alike, choose the clearest fit and stop unless a new idea could change the decision.
- When evidence is thin, keep the honest default. Revisit if enough new traffic or a distinct new angle makes another test useful.
- If viewers leave because the content breaks the promise, fix that mismatch before polishing more words.

Do not keep testing until a preferred option wins. Do not use hourly swings to call a result.

Return the decision, supporting report values, limits, and one next action. Add a log row: episode, dates, question, option IDs, tool result, decision, and lesson status. Mark the lesson as a guess, a result for this episode, or a repeated pattern.
