# Shadowing Player AI Production Recipe

- Recipe version: 1.1.1
- LessonDocument version: 1 (`"schema": "shadowing-player-lesson"`, `"version": 1`)
- Last updated: 2026-10-09

This file is a production specification. Give it to an AI assistant (or a person) together with what you want to teach, and it describes exactly how to plan the script, the video, the subtitles and the lesson file so that the result works in Shadowing Player.

Shadowing Player is a browser extension that lets a learner practise a YouTube video one sentence at a time: repeat it, shadow it, recall it from its translation, or answer a quiz over the video. The video stays on YouTube. A *lesson* is a small text file that says where each sentence is in the video and what it means.

## 0. Rules that override everything else

1. **Compatibility is more important than creativity.** Use only the fields, modes and behaviours described in this file. Never invent a field, a mode, a schema property or a timing rule. If something the author asks for cannot be expressed with what is described here, say so plainly and offer the closest thing that can.
2. **Do not create final timestamps before the video edit is locked.** A lesson is a list of times in one specific video file. Any trim, re-order or re-encode after timing shifts every sentence after it. Draft times are fine for planning; final times are written only for the final, uploaded video.
3. **Never produce times you cannot know.** An AI that has not been given the finished video or its measured timings cannot know them. In that case produce the script, the cue table *without* final times (or with clearly labelled planned durations), and instructions for timing. Do not present guessed times as real.
4. **Sentence IDs are permanent.** Once a lesson is published, an ID is never renamed, renumbered or reused.
5. **Only material the author has the right to use.** Do not copy subtitles, captions or scripts from other people's videos. Do not fetch or scrape captions from YouTube.
6. **The video works on its own; the player makes it richer.** Someone watching the plain YouTube video, with no extension, must be able to follow it from start to finish: nothing is missing, nothing refers to a button they cannot see, and the story or lesson makes sense in order. Shadowing Player then adds what a video cannot do by itself: pausing, repeating, recall, choosing an answer, being told right or wrong, practising one unit. Design the video first as a good video, then describe it in the lesson so the player can do more with it. Never make a video that only makes sense inside the player.
7. **The video never draws the quiz.** Answer choices, A/B/C/D buttons, "choose the answer" screens, countdown timers, progress bars for thinking time, scores and right/wrong marks are drawn by the player over the paused video. None of them is rendered into the video. The video contains only the question and the answer (section 7, Quiz).
8. **Learner state is not authored.** Favorites, mastered marks, practice counts and scores belong to each learner and live in their browser. They are never written into a lesson.

## 1. What Shadowing Player does with a lesson

Everything the player does comes down to one action: **seek to the `start` of a sentence, play to its `end`, stop.** Every practice mode is built from that.

| Mode | What the learner experiences | What it needs from the lesson |
| --- | --- | --- |
| Normal | Ordinary playback, sentence by sentence. Previous / Next. Optional Repeat: 1, 2, 3, 5 times or until stopped | Sentences with clean boundaries |
| Shadowing | One sentence plays, the video pauses so the learner can say it, then it repeats or moves on. Repeat N means N rounds of listen-then-speak | Sentences with clean boundaries |
| Recall | The sentence is shown (and optionally read aloud by the learner's browser voice) in the language the learner knows; a short countdown; then the video plays the sentence as the answer | Every sentence has both `en` and `ko` text; `videoLanguage` is set |
| Quiz | A sentence plays, the video pauses, a multiple-choice question appears over the video; the chosen answer can jump to its own explanation sentence | `quizzes`, with a trigger sentence and usually one feedback sentence per choice |

Things that are *options*, not modes:

- **Order**: Linear (video order) or Random (shuffled; every sentence of the chosen scope comes up once before any repeats). Random applies to Shadowing and Recall, and to Normal when Repeat is on. A quiz always follows the order of `quizzes`.
- **Target scope**: the learner practises all sentences, their own favorites, or one of the lesson's `groups` (section 8).
- **Skip mastered**: the learner can leave out sentences they marked as mastered. It does not apply to Quiz.
- **Speaking time**: created by the player (a fixed 0, 1, 2, 3 or 5 seconds, or a time that follows the length of the sentence). It is never recorded into the video.
- **Speed**: 0.5× to 1.5×, chosen by the learner.

The player supports **two languages, English and Korean**. A sentence has an `en` text and a `ko` text. No other language field exists.

## 2. Questions to settle before producing anything

Determine each of these. If the author did not say, ask briefly, or choose the conservative default in brackets and state that you did.

| Question | Default |
| --- | --- |
| Which language are the learners learning (the *target*)? English or Korean | (must be known) |
| Which language do they already know (the *source*)? | the other one |
| Teaching goal and topic | (must be known) |
| Learner level | beginner to lower-intermediate |
| Number of practice sentences | 30 |
| Video format: target language only, or each sentence said in both languages | target language only |
| If both: which language is said first | source first, then target |
| Which practice modes matter | Shadowing and Recall |
| Grouping strategy | by the natural structure of the content (section 8) |
| Quiz wanted? | no |
| Branching feedback for quiz answers? | yes, if a quiz is wanted |
| Is the video already made, or still to be made? | still to be made |

## 3. The lesson file (LessonDocument version 1)

A JSON object. These are **all** the fields that exist. Anything else is ignored by the player and must not be written.

### Root

| Field | Required | Rule |
| --- | --- | --- |
| `schema` | yes | Exactly `"shadowing-player-lesson"` |
| `version` | yes | The number `1` |
| `cues` | yes | 1 to 20,000 sentences |
| `videoId` | no | Text. The YouTube video ID. Informational |
| `title` | no | Text |
| `language` | no | Text, informational (for example `"en-ko"`) |
| `videoLanguage` | no, but set it | `"en"` or `"ko"`: the language spoken in each sentence's own `start`–`end` span. Recall needs it |
| `groups` | no | Up to 500. Section 8 |
| `quizzes` | no | Up to 2,000. Section 9 |
| `practice` | no | `{ "defaultOrder": "linear" | "random" }`. Stored but **not applied**. Do not rely on it |
| `audioPrompts` | no | **Reserved and unused.** Do not write it |

### A sentence (`cues[]`)

| Field | Required | Rule |
| --- | --- | --- |
| `id` | yes | 1 to 64 characters from `A–Z a–z 0–9 _ . -`. Unique among the sentences. Permanent once published |
| `start` | yes | Seconds from the start of the video, a decimal number, 0 or more |
| `end` | yes | Seconds, at least 0.05 after `start` |
| `en` | at least one of the two | English text, up to 2,000 characters |
| `ko` | at least one of the two | Korean text, up to 2,000 characters |
| `note` | no | Optional text |
| `prompt` | no | Section 6. A timed span, **not text** |

Give every sentence **both** `en` and `ko`. A sentence with only one cannot be used in Recall.

Sentences must not overlap each other in time.

### What does not exist

There is no field for: audio or video data, image data, a URL of any kind, a TTS provider or voice, favorites, mastered state, scores, difficulty, tags, speaker names, pauses or speaking time, playback speed, repeat count, a third language, per-sentence instructions, or timestamps for a quiz branch. Do not add them under any name.

## 4. Designing sentences

| Property | Target |
| --- | --- |
| Length of one sentence | **3 to 5 seconds** is best. 2 to 7 seconds is acceptable. Never longer than 7: split at a natural pause |
| Content | One sentence, one thought |
| Quiet before the first sound (pre-roll) | 150 to 300 ms |
| Quiet after the last sound (post-roll) | 250 to 500 ms |
| Gap between two sentences | 300 to 800 ms: simply the post-roll of one plus the pre-roll of the next |
| Shortest useful sentence | about 0.6 seconds of speech. One-word lines ("Yes.") are poor practice items |

- `start` sits inside the pre-roll, just before the first sound (at least 0.05 s before it). `end` sits inside the post-roll, just after the last sound.
- **Do not record silence for the learner to speak in.** The player makes the speaking time. Five seconds of silence after each sentence only makes the video long and the lesson no better.
- Each sentence should be worth repeating: natural, complete, at the learner's level.
- Write the translation the way it would be *said*. No notes in brackets ("(formal)"), no alternatives with slashes ("go / leave"). In Recall the translation is read aloud and is the question; it must mean the same as the sentence, at the same level of politeness, and suggest one clear answer.

### Stable IDs

- Use zero-padded IDs that leave room to grow: `c001`, `c002`, … for up to 999 sentences; `c0001` for more.
- A new sentence added later gets a **new** ID (`c121`), never the ID of a removed one, and existing IDs are not shifted to make room.
- Changing a sentence's text or times is safe. Changing its `id` detaches every learner's progress from it and breaks quiz branches that point at it.
- Quiz IDs are permanent in the same way.

## 5. Recording and editing

- **Speech dominates.** No background music during practice sections, or music so low it is barely there.
- **Nothing crosses a sentence boundary**: no music bed, no reverb tail, no transition sound, no sound effect, no second speaker starting early. A cut inside any of these is heard every time the sentence repeats.
- **One speaker at a time**, close and dry microphone sound.
- **Do not cut sentences tight against each other.** Keep the handles of quiet.
- Keep a meaningful picture at the start of each sentence: the player seeks there, so that frame is what the learner sees first.
- Leave at least half a second after the last sentence before the video ends.
- An intro of 10 to 20 seconds is fine. It is not part of the lesson's sentences.
- **Lock the edit, export, upload. Only then write final times**, measured on the final file.

Rough length: 100 sentences of 3 to 5 seconds, said once each, make a video of about 8 to 12 minutes.

## 6. Videos that say each sentence in two languages (`cues[].prompt`)

A video may say every sentence twice: first in the language the learner knows, then in the language they are learning (or the other way round). This is supported, and the field that describes it is called `prompt`.

**What `prompt` is: a span of time in the video and the language spoken in it.**

```json
{ "id": "c002", "start": 12.6, "end": 15.0,
  "en": "Can I get it to go?", "ko": "포장해 주실 수 있나요?",
  "prompt": { "start": 10.75, "end": 12.25, "lang": "ko" } }
```

| Field | Rule |
| --- | --- |
| `prompt.start` | Seconds, 0 or more |
| `prompt.end` | Seconds, at least 0.05 after `prompt.start` |
| `prompt.lang` | `"en"` or `"ko"`: the language spoken in that span |

- The sentence's own `start`–`end` is where it is spoken in the language named by `videoLanguage` (the *answer*). `prompt` is where the same sentence is spoken in the other language.
- The words of the prompt are simply the sentence's own `ko` or `en` text. `prompt` holds no text.

**What `prompt` is not.** It is not the question text of Recall. It is not text for a voice to read. It is not a quiz question. It is not a reference to another sentence. It carries no audio. Recall and Quiz never use it.

**What the learner gets from it** (in Normal and Shadowing): a choice of what one "sentence" plays (both parts in video order, the answer only, or the prompt only); an optional pause between the two parts (none, 1 second, 3 seconds, or a 3-2-1 count over the video during which the answer's text is hidden); and a separate playback speed for each language.

Production rules for such a video:

- Give a `prompt` to **every** sentence or to none.
- Keep clean handles around **both** parts, and do not let the two spans overlap.
- Leave a natural gap of about 300 to 500 ms between the two parts. **Do not record thinking time between them**: the player adds the pause the learner chooses.
- Use the same language first throughout the video.
- Set `videoLanguage` to the language of the *second-role* part, the one in `start`–`end` (for learners of English: `"en"`).

A video that speaks only the target language needs no `prompt` at all, and Recall still works: the player shows and reads the translation itself.

## 7. Recipes for each way of practising

### Shadowing

- Natural speech at real speed; one meaningful phrase per sentence; 2 to 7 seconds.
- Clean start and clean end are what matter. The learner will hear each boundary many times.
- Show the speaker's face and mouth when pronunciation is the point.
- No silence for the learner in the video.

### Recall

- The learner sees the sentence in the language they know, says it in the language they are learning, then the video plays the sentence. The video is the answer key.
- Needs: both texts on every sentence, and `videoLanguage`. Nothing else.
- For learners of English: the video speaks English, `videoLanguage` is `"en"`, the question shown is the `ko` text. For learners of Korean: the mirror image.
- The first syllable of every sentence must be clean: the learner has just spoken and is listening for it.
- Make one video per target language. If `videoLanguage` does not match what the learner is learning, the answer is only shown as text and the video is not played.
- The countdown and the speaking time are the player's. Do not record them.

### Random

- In Random order a sentence is heard with nothing before it. Use sentences that stand on their own: "Where are you going?" works; "That one." does not.
- Keep dialogue that depends on its neighbours in its own group, so that learners shuffle the sentence bank and play the dialogue in order.

### Quiz

**Read this before designing any quiz video. The most common mistake is to render the multiple-choice screen into the video. Do not.**

#### Who does what

| Thing | Where it lives | Who shows it |
| --- | --- | --- |
| The question (a sentence, a situation, a line of dialogue) | **In the video**, as a sentence (a cue) | The video |
| The answer and its explanation | **In the video**, as another sentence (a cue) | The video |
| The question text on the card | In the lesson file: `quizzes[].prompt` | The player, over the paused video |
| The answer choices (A, B, C, D) | In the lesson file: `quizzes[].choices[].text` | The player, as buttons over the paused video |
| Which choice is right | In the lesson file: `choices[].correct` | The player, as "Correct" / "Not quite" |
| The pause to think and choose | **Nowhere.** The player pauses the video and waits as long as the learner needs | The player |
| Timers, countdown bars, score, question counter | **Nowhere.** Not in the video, not in the lesson | The player shows its own progress and score |

So a quiz video is simple. **For each question the video contains two things, one after the other: the question, then the answer.** Nothing else.

This is also what makes the video good on its own (section 0, rule 6). A person watching on YouTube without the player hears the question, thinks for a moment, and hears the answer: a complete quiz video. A person using the player gets more from the same video: it stops and waits for them, offers choices, tells them whether they were right, and keeps score.

```
VIDEO      [ question sentence ]   [ answer sentence: the right answer, and why ]
                               ↑
PLAYER     pauses here on the last frame of the question,
           draws the question text and the choice buttons from the lesson file,
           waits for the learner to pick one,
           shows "Correct" or "Not quite",
           then plays the answer sentence (or the sentence that choice points to),
           then stops, or goes on to the next question (the learner's setting).
```

#### The frames of the video

- **During the question**: show and say the question only. For "Which is the most natural English for this?" that means the Korean sentence on screen, and nothing that looks like a set of options.
- **Never on screen**: the choices, letters A/B/C/D, boxes that look like buttons, "choose the best answer", a timer or a shrinking bar, a score.
- **The player's card covers the middle of the paused frame** (roughly from 11% below the top to 19% above the bottom), and dims the picture slightly. Put the question text in the lesson's `quizzes[].prompt` as well, because the card will sit over the middle of the frame. Keep titles and anything else that must stay visible near the top or bottom edge.
- **End the question on a steady frame.** Not on a cut, a fade or a blur: that frame is what stays on screen while the learner chooses.
- **During the answer**: say the right answer, and show it if you like. Say it so that it is true whether the learner was right or wrong: "The natural way to say it is 'I'm just window shopping.' 'Eye shopping' is Konglish." Do not open with "Correct!" or "Wrong!": the player has already said which.
- **A short beat to think, not a waiting screen.** Leave one to three seconds between the question and the answer, so the video also works for someone watching without the player. Put it *between* the two sentences: it belongs to neither cue, so the player skips it. Do not record a long silence, a timer or a "choose now" screen: in the player the pause lasts as long as the learner needs, and on plain YouTube a viewer can pause for themselves.
- **Nothing in the video refers to the player.** No "tap your answer", "pick A, B, C or D", "click below". Say "What would you say?" or "Think about it", which is true for every viewer.
- Clean handles around both sentences, as everywhere (section 4).

#### The simple quiz (use this unless asked for more)

Two sentences per question: the question (`q01`) and the answer (`q01a`). Every choice leads to the same answer sentence.

```json
"cues": [
  { "id": "q01",  "start": 12.0, "end": 15.2, "en": "", "ko": "나 그냥 아이쇼핑 하고 있어.", "note": "question: the Korean sentence is shown and said. No English here: it would be the answer" },
  { "id": "q01a", "start": 16.0, "end": 21.5, "en": "The natural way to say it is \"I'm just window shopping.\" \"Eye shopping\" is Konglish.", "ko": "자연스러운 표현은 \"I'm just window shopping.\"이에요. \"Eye shopping\"은 콩글리시예요.", "note": "answer and explanation" }
],
"quizzes": [
  {
    "id": "konglish-01",
    "triggerCueId": "q01",
    "prompt": "나 그냥 아이쇼핑 하고 있어. — 가장 자연스러운 영어는?",
    "choices": [
      { "id": "a", "text": "I'm just eye shopping.",       "correct": false, "targetCueId": "q01a" },
      { "id": "b", "text": "I'm just window shopping.",    "correct": true,  "targetCueId": "q01a" },
      { "id": "c", "text": "I just shopping with eyes.",   "correct": false, "targetCueId": "q01a" },
      { "id": "d", "text": "I'm just looking shopping.",   "correct": false, "targetCueId": "q01a" }
    ]
  }
]
```

What the learner experiences:

1. The video plays `q01` and pauses on its last frame.
2. The player shows the question text and the four choices over the video.
3. The learner picks one. The player shows "Correct" or "Not quite".
4. The player plays `q01a`, so the learner hears the right answer and the reason either way, and pauses.
5. The player stops there, or moves to the next question, according to the learner's own setting.

**The question sentence must not contain the answer.** While the question plays, the player shows that sentence's `en` and `ko` in its panel, over the video and in the list of sentences, like any other sentence. So a question sentence holds only what the question says. In a "how do you say this in English?" quiz the question is the Korean sentence: give it `ko` and leave `en` empty (`""`), as above. Do not fill `en` with the translation, because the translation is the answer. A sentence needs text in at least one of the two languages, not both. The same applies the other way round, and to a question whose answer is a single word: nothing in the question sentence's text may give it away.

A choice may also leave `targetCueId` out. Then, after the verdict, nothing is played for that choice. Use this only when there is nothing useful to hear.

A ten-question quiz video is therefore twenty sentences: `q01`, `q01a`, `q02`, `q02a`, … Quizzes are asked in the order they are written in `quizzes`.

#### Branching (optional: a different explanation for each choice)

When the author is willing to record more, each choice can lead to its **own** explanation sentence instead of a shared one. This teaches more, because a wrong answer is explained for what it is.

- One feedback sentence per choice, each a cue with its own ID (`q01a`, `q01b`, `q01c`, …), each explaining *that* choice.
- **A branch names a sentence (`targetCueId`), never a time.** There is no way to branch to a raw timestamp, and none should be invented.
- Every feedback sentence is complete in itself and has clean handles. It is heard alone, so it must not say "as I just said".
- The player jumps to the chosen sentence, plays to its `end` and pauses. It never plays on into the next one.
- Order on the timeline is free, but the video must still read well straight through (section 0, rule 6). A layout that does: question, then the explanations one after another as a walk-through of the options: "'Eye shopping' is what many people say first, but it is Konglish. … The natural phrase is 'window shopping'." Each sentence names the option it is about, so it makes sense both in the walk-through and when the player jumps to it alone. Do not write them as reactions ("You chose A. Wrong!"), which make no sense to someone who did not choose.

The worked example in section 15 uses branching.

#### For any quiz

- Quizzes check understanding at useful points. They are not needed on every sentence.
- Two to twelve choices; three or four work best. One line each. The question in one or two lines.
- Exactly one choice is correct. Vary its position from question to question: over a whole lesson each position (first, second, third, fourth) should be the right one about equally often, with no pattern.
- Write the question in the language learners read comfortably, and the choices in the language being taught.

## 8. Groups: giving a long lesson useful parts

A lesson with hundreds of sentences should not be one flat list. `groups` lets the learner choose a part of the lesson as their **target scope** and practise only that.

```json
"groups": [
  { "id": "present-simple", "title": "현재단순 · Present Simple", "cueIds": ["c001", "c002", "c003", "c004"] }
]
```

| Field | Rule |
| --- | --- |
| `id` | 1 to 64 characters from `A–Z a–z 0–9 _ . -`. Unique among groups. Permanent once published. **Not shown to learners** |
| `title` | What learners read, exactly as written. Any language. Not empty |
| `cueIds` | At least one existing sentence ID. No ID twice in the same group |

- **A sentence may belong to several groups, or to none.** One sentence can be in "Past Simple", in "SHE" and in "Negative" at once.
- Groups are offered in the order written. Inside a group, sentences are practised in video order.
- At most 500 groups.
- A quiz belongs to a scope when its trigger sentence is in it.

Useful ways to group, alone or combined:

| By | Examples |
| --- | --- |
| Grammar | Present Simple, Past Simple, Past Progressive |
| Person | I, YOU, HE, SHE, IT, WE, THEY |
| Sentence form | Affirmative, Negative, Question |
| Topic | Airport, Restaurant, Hotel |
| Difficulty | Beginner, Intermediate |
| Chapter | Chapter 1, Chapter 2 |

Size guidance:

| Lesson size | Grouping |
| --- | --- |
| About 100 sentences | 4 to 10 groups of 10 to 25 sentences |
| About 300 sentences | Two dimensions, for example chapter and form. A learner should be able to reach a set of 20 to 60 sentences |
| 600 or more | Three dimensions (for example tense × person × form). Every group should still be a session someone would actually sit down to: roughly 30 to 120 sentences |

Learners also always have **All**, and **Favorites** once they star a sentence. Do not create groups called "favorites" or "mastered": those are the learner's own.

## 9. Quizzes in the lesson file

```json
"quizzes": [
  {
    "id": "to-go-1",
    "triggerCueId": "q01",
    "prompt": "Which one is natural?",
    "choices": [
      { "id": "a", "text": "I go with coffee.",   "correct": false, "targetCueId": "q01b" },
      { "id": "b", "text": "Can I get it to go?", "correct": true,  "targetCueId": "q01a" },
      { "id": "c", "text": "Give me go.",         "correct": false, "targetCueId": "q01c" }
    ]
  }
]
```

| Field | Required | Rule |
| --- | --- | --- |
| `id` | yes | Same character rule as other IDs. Unique among quizzes. Permanent |
| `triggerCueId` | yes | An existing sentence ID |
| `prompt` | yes | The question text. (Same word as `cues[].prompt`, different thing: here it **is** text.) |
| `choices` | yes | 2 to 12. Three or four work best |
| `choices[].id` | yes | Unique within the quiz |
| `choices[].text` | yes | Shown on the button exactly as written |
| `choices[].targetCueId` | no | An existing sentence ID: the feedback sentence for this choice. Without it only the verdict is shown |
| `choices[].correct` | no | `true` or `false` |

- If any choice has `correct`, **exactly one** must be `true`.
- If no choice has `correct`, the question only branches: nothing is marked right or wrong.
- Questions are asked in the order of `quizzes`.
- Write the question in the language learners read comfortably and the choices in the language being taught.

## 10. Subtitle files and what each format can carry

| Output | Carries | Does not carry |
| --- | --- | --- |
| **LessonDocument JSON** | Everything in this file | — |
| **SRT / VTT** | Sentences with times and text. With an English line and a Korean line in each caption, both languages | Stable IDs, `videoLanguage`, `prompt`, groups, quizzes |

A subtitle file is enough for Normal, Shadowing, and (with both lines) Recall. Anything involving groups, quizzes, branching or two-language timing needs the JSON.

**Two subtitle files, one English and one Korean**, can be uploaded together in Teacher Studio. They are joined by rule, not by meaning: when both files have the same number of captions they are paired in order; otherwise captions that overlap in time are put together. Nothing is translated. The author must check the pairs.

**Automatic captions** (for example YouTube's) may be used as a starting point by the owner of the video, who downloads them as a file. They contain recognition and timing errors and are cut into fragments, not sentences. They must be reviewed and re-cut by a person before publishing. Do not fetch captions from YouTube by any other means. Importing captions directly from YouTube inside Teacher Studio is not available yet.

## 11. Speech and voices

- The voice that reads a Recall question, and any spoken guidance, is the player's and the learner's browser's concern. A lesson contains no audio and names no voice.
- Do not put synthesized readings of the translation into the video "for Recall". If the lesson design is a two-language video, that is section 6, recorded like any other speech.

## 12. Testing before publishing

A production plan must end with these checks, done by a person in Shadowing Player on the published video.

| Check | How | Pass when |
| --- | --- | --- |
| Three speeds | Play several sentences at 0.75×, 1.0×, 1.25× | No clipped start or end at any speed |
| Three repeats | Repeat one sentence 3 times | Same clean boundary each time; nothing of the next sentence leaks in |
| First and last syllable | Listen to the first and last sound of at least ten sentences | Both whole |
| Random jumps | Random on, step through ten sentences | Each starts cleanly after a jump from far away, and makes sense alone |
| Recall | Twenty sentences in a row | Question and answer mean the same; the answer starts cleanly |
| Quiz | Every question | Pauses at the right frame; the card does not cover what matters; **no choices, timer or score are visible in the video itself** |
| Branching | **Every choice of every question** | Goes to its own sentence, plays all of it and nothing more, pauses |
| Two-language transitions | If the video has `prompt` spans: both, answer only, prompt only; pause none and 3-2-1 | Each part starts and ends cleanly; the parts do not bleed into each other |
| Groups | Choose each group once | It contains what its title says |

## 13. Do not do this

- Five seconds of silence after every sentence.
- One 20-second cue holding several sentences.
- Music, reverb or a transition sound across a sentence boundary.
- IDs that change when sentences are re-ordered or re-timed; renumbering after an edit; reusing an old ID.
- Branching to a timestamp, or to "about 3:20".
- Inventing LessonDocument fields ("difficulty", "audioUrl", "pauseAfter", "speaker", "mastered": true).
- Putting text into `cues[].prompt`, or treating it as the Recall question.
- Writing `favorite` or `mastered` into a lesson.
- Publishing raw automatic captions without review.
- Using subtitle packs, scripts or captions you have no right to use.
- Final timestamps made before the edit was locked, or times guessed without the video.
- A group so large it is the whole lesson again, or dozens of groups of three sentences.
- A quiz on every sentence.
- Rendering the answer choices, A/B/C/D buttons, a countdown bar or a score into the video. The player draws these; the video only asks and answers.
- Recording a long silence, or a timer, for the learner to choose in.
- Putting the answer into the question sentence's own text, for example its translation in `en`. The player shows that text while the question plays.
- A video that only makes sense inside the player: lines such as "tap your answer", or explanations written as reactions to a choice the viewer never made.
- An answer sentence that begins "Correct!" when wrong answers lead to it too.
- One feedback sentence shared by all wrong answers.

## 14. What to return

When asked to plan a lesson, return these parts, in this order. Leave out a part only when it does not apply, and say that it does not.

1. **Lesson concept**: target and source language, level, goal, video format, modes it is designed for.
2. **Script**: every sentence in both languages, in recording order.
3. **Cue table**: `id`, text in both languages, planned duration; real `start` / `end` only if measured from the final video.
4. **Group design**: each group's `id`, `title` and sentence IDs, and why it is a useful session.
5. **Quiz design**: for each question, the question sentence, the card text, the choices and which is correct. State explicitly that the choices are lesson data and are not rendered into the video.
6. **Branching design**: the answer sentence each choice leads to (one shared answer sentence in the simple quiz, or one per choice).
7. **Video editing instructions**: order of material, what is on screen in each part, handles, what must not cross boundaries, and for a quiz, a frame-by-frame list of what the video shows during the question and during the answer.
8. **Subtitle export plan**: which files will be made and from what.
9. **LessonDocument plan**: or the JSON itself, if asked and if the times are known.
10. **Human QA checklist**: section 12, adapted to this lesson.

**If asked for the LessonDocument JSON itself**: output one valid JSON object that uses only the fields of section 3, 8 and 9, with every ID rule obeyed and every reference pointing at an existing sentence. If final times are not known, say so and do not output invented times as if they were final.

## 15. Worked example

A short lesson for learners of English whose first language is Korean. The video says each sentence in Korean, then in English. Eight practice sentences in two tenses, grouped two ways, and one quiz question with a feedback sentence for each choice.

Its times are an example and match no real video.

```json
{
  "schema": "shadowing-player-lesson",
  "version": 1,
  "title": "At the café: present and past",
  "language": "en-ko",
  "videoLanguage": "en",
  "cues": [
    {"id": "c001", "start": 7.85, "end": 10.25, "en": "I'd like an iced americano.", "ko": "아이스 아메리카노 한 잔 주세요.", "prompt": {"start": 6.0, "end": 7.5, "lang": "ko"}},
    {"id": "c002", "start": 12.6, "end": 15.0, "en": "Can I get it to go?", "ko": "포장해 주실 수 있나요?", "prompt": {"start": 10.75, "end": 12.25, "lang": "ko"}},
    {"id": "c003", "start": 17.35, "end": 19.75, "en": "I don't need a receipt.", "ko": "영수증은 필요 없어요.", "prompt": {"start": 15.5, "end": 17.0, "lang": "ko"}},
    {"id": "c004", "start": 22.1, "end": 24.5, "en": "Do you have oat milk?", "ko": "귀리 우유 있나요?", "prompt": {"start": 20.25, "end": 21.75, "lang": "ko"}},
    {"id": "c005", "start": 26.85, "end": 29.25, "en": "She ordered a latte.", "ko": "그녀는 라테를 주문했어요.", "prompt": {"start": 25.0, "end": 26.5, "lang": "ko"}},
    {"id": "c006", "start": 31.6, "end": 34.0, "en": "She didn't want sugar.", "ko": "그녀는 설탕을 원하지 않았어요.", "prompt": {"start": 29.75, "end": 31.25, "lang": "ko"}},
    {"id": "c007", "start": 36.35, "end": 38.75, "en": "Did she pay by card?", "ko": "그녀는 카드로 결제했나요?", "prompt": {"start": 34.5, "end": 36.0, "lang": "ko"}},
    {"id": "c008", "start": 41.1, "end": 43.5, "en": "We sat by the window.", "ko": "우리는 창가에 앉았어요.", "prompt": {"start": 39.25, "end": 40.75, "lang": "ko"}},
    {"id": "q01", "start": 48.5, "end": 51.9, "en": "You want to take your coffee with you. What do you say?", "ko": "커피를 가지고 나가고 싶어요. 뭐라고 말할까요?", "prompt": {"start": 45.25, "end": 48.15, "lang": "ko"}},
    {"id": "q01a", "start": 56.45, "end": 60.45, "en": "The natural way to ask is \"Can I get it to go?\"", "ko": "자연스러운 표현은 \"Can I get it to go?\"예요.", "prompt": {"start": 52.5, "end": 56.1, "lang": "ko"}},
    {"id": "q01b", "start": 66.0, "end": 71.4, "en": "\"I go with coffee\" says that you are leaving, not what you are asking for.", "ko": "\"I go with coffee\"는 내가 떠난다는 말이지, 부탁하는 말이 아니에요.", "prompt": {"start": 61.05, "end": 65.65, "lang": "ko"}},
    {"id": "q01c", "start": 76.35, "end": 80.95, "en": "\"Give me go\" is not a sentence, and it sounds like an order.", "ko": "\"Give me go\"는 문장이 아니고, 명령처럼 들려요.", "prompt": {"start": 72.0, "end": 76.0, "lang": "ko"}}
  ],
  "groups": [
    {"id": "present-simple", "title": "현재단순 · Present Simple", "cueIds": ["c001", "c002", "c003", "c004"]},
    {"id": "past-simple", "title": "과거단순 · Past Simple", "cueIds": ["c005", "c006", "c007", "c008"]},
    {"id": "affirmative", "title": "긍정 · Affirmative", "cueIds": ["c001", "c005", "c008"]},
    {"id": "negative", "title": "부정 · Negative", "cueIds": ["c003", "c006"]},
    {"id": "question", "title": "의문 · Question", "cueIds": ["c002", "c004", "c007"]}
  ],
  "quizzes": [
    {
      "id": "to-go-1",
      "triggerCueId": "q01",
      "prompt": "Which one is natural?",
      "choices": [
        {"id": "a", "text": "I go with coffee.", "correct": false, "targetCueId": "q01b"},
        {"id": "b", "text": "Can I get it to go?", "correct": true, "targetCueId": "q01a"},
        {"id": "c", "text": "Give me go.", "correct": false, "targetCueId": "q01c"}
      ]
    }
  ]
}
```

What it shows:

- `videoLanguage` is `"en"`: each sentence's own span is the English. Every sentence has a `prompt` span with `"lang": "ko"`: where the Korean is spoken, about 0.35 s before the English starts.
- IDs are stable and zero-padded. Quiz sentences have their own series (`q01`, `q01a`, …).
- Five groups over eight sentences: `c006` ("She didn't want sugar.") is in both "Past Simple" and "Negative". The quiz sentences belong to no group and are still reachable under All.
- Group IDs are plain ASCII; titles are written for learners, in two languages.
- The quiz has exactly one correct choice, the right answer is not first, and each choice leads to its own feedback sentence by ID. This is the branching form; the simple form, with one shared answer sentence, is in section 7.
- Nothing about favorites, mastery, pauses, speeds or voices appears anywhere.

## 16. Prompt to copy

Attach this file, then send:

```text
I am creating a Shadowing Player lesson.
Follow the attached Shadowing Player AI Production Recipe strictly.
Do not invent fields, modes or behaviours that the recipe does not describe.
If something I ask for cannot be done within the recipe, tell me.

Topic: [TOPIC]
Learners are learning: [English / Korean]
Learners already know: [Korean / English]
Level: [LEVEL]
Number of practice sentences: [NUMBER]
Video format: [target language only / each sentence in both languages, which first]
Practice modes that matter: [Shadowing / Recall / Random / Quiz]
Quiz: [none / how many questions; simple (question then answer) or branching (an explanation per choice)]
The video is: [not made yet / finished, and I will give you the timings]

Return the ten parts listed under "What to return".
Do not output final timestamps unless I have given you timings measured on the finished video.
The video must be complete and enjoyable for someone watching it on YouTube without Shadowing Player; the player adds interaction on top.
If there is a quiz: the video shows only the question and then the answer. Do not design any screen that shows the answer choices, a timer or a score; those come from the lesson file and are drawn by the player.
```
