For agentic workers: REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (
- [ ]) syntax for tracking.
Goal: Replace the dual TTS owners in DialogueChatView with a single view-level useAudioPlayer composable backed by Azure Speech REST, enforce single-channel + recording-interrupt rules structurally, and surface playback errors per message with one-click retry.
Architecture: A new useAudioPlayer composable owns the only HTMLAudioElement and the only synthesis AbortController, exposing play(id, source) / stop() / three reactive ids (playingId, loadingId, errorId). A new stateless speechService.synthesize() calls Azure Speech REST and returns an MP3 blob. useDialogueEngine is stripped of all speechSynthesis references and gains attachStudentBlob() so the WebSocket path can attach the recorded blob to the student message it created. DialogueChatView watches for AI messages reaching done and triggers auto-play; recording start calls player.stop(); the play button on every voice-bar renders one of four states (idle/loading/playing/error) directly from the player refs.
Tech Stack: Vue 3 <script setup> + TypeScript, Vite (import.meta.env), browser fetch + AbortController, browser HTMLAudioElement, Azure Speech REST API.
Repo root: /Users/buoy/Development/gitrepo/PPT
src/views/Editor/EnglishSpeaking/services/speechService.ts — stateless Azure REST synthesis. One exported function synthesize(text, signal): Promise<Blob>.src/views/Editor/EnglishSpeaking/composables/useAudioPlayer.ts — sole audio owner. Cache, abort, single-channel rule, error surface..env.example (or append if exists) — document the two new env vars.src/views/Editor/EnglishSpeaking/composables/useDialogueEngine.ts — remove speakTTS / cancelTTS / ttsUtterance and all 4 call sites; add attachStudentBlob.src/views/Editor/EnglishSpeaking/preview/DialogueChatView.vue — instantiate useAudioPlayer, replace local currentAudio/stopCurrentPlayback/playingMessageId, rewrite togglePlay, fix WS path's missing audioBlob, add auto-play watcher, call player.stop() from handleStartRecording / handleRestart, render the 4-state play buttons + error hint.Spec reference: docs/superpowers/specs/2026-04-26-azure-tts-audio-player-design.md
Package manager: npm (project has package-lock.json).
Adds the stateless synthesis function and documents required environment variables. Self-contained — nothing imports it yet.
Files:
/Users/buoy/Development/gitrepo/PPT/src/views/Editor/EnglishSpeaking/services/speechService.tsCreate or append: /Users/buoy/Development/gitrepo/PPT/.env.example
[ ] Step 1: Create the service file
Create src/views/Editor/EnglishSpeaking/services/speechService.ts:
const KEY = import.meta.env.VITE_AZURE_SPEECH_KEY as string | undefined
const REGION = import.meta.env.VITE_AZURE_SPEECH_REGION as string | undefined
const VOICE = 'en-US-AriaNeural'
const FORMAT = 'audio-24khz-48kbitrate-mono-mp3'
/**
* Synthesize English text via Azure Speech REST.
* Returns an MP3 Blob. Throws on credential / network / non-2xx.
*
* Pass an AbortSignal so callers (the audio player) can cancel
* an in-flight synthesis when the user starts recording or
* triggers a different playback.
*/
export async function synthesize(text: string, signal?: AbortSignal): Promise<Blob> {
if (!KEY || !REGION) {
throw new Error('Azure Speech credentials not configured (VITE_AZURE_SPEECH_KEY / VITE_AZURE_SPEECH_REGION)')
}
const ssml =
`<speak version='1.0' xml:lang='en-US'>` +
`<voice name='${VOICE}'>${escapeXml(text)}</voice>` +
`</speak>`
const res = await fetch(
`https://${REGION}.tts.speech.microsoft.com/cognitiveservices/v1`,
{
method: 'POST',
signal,
headers: {
'Ocp-Apim-Subscription-Key': KEY,
'Content-Type': 'application/ssml+xml',
'X-Microsoft-OutputFormat': FORMAT,
'User-Agent': 'PPT-EnglishSpeaking',
},
body: ssml,
},
)
if (!res.ok) {
throw new Error(`Azure TTS failed: ${res.status} ${res.statusText}`)
}
return res.blob()
}
function escapeXml(s: string): string {
return s.replace(/[<>&'"]/g, c => ({
'<': '<',
'>': '>',
'&': '&',
"'": ''',
'"': '"',
}[c]!))
}
If .env.example does not exist, create it with the contents below. If it exists, append the two new lines (after a blank line).
VITE_AZURE_SPEECH_KEY=
VITE_AZURE_SPEECH_REGION=
Run:
cd /Users/buoy/Development/gitrepo/PPT
npm run type-check
Expected: exits 0, no errors.
cd /Users/buoy/Development/gitrepo/PPT
git add src/views/Editor/EnglishSpeaking/services/speechService.ts .env.example
git commit -m "$(cat <<'EOF'
feat: add Azure Speech REST synthesis service
Stateless synthesize(text, signal) returning an MP3 Blob via
the Azure Speech REST API. Reads VITE_AZURE_SPEECH_KEY and
VITE_AZURE_SPEECH_REGION; documents both in .env.example.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
EOF
)"
The sole audio playback owner. Owns the only HTMLAudioElement, the only synthesis abort controller, and the cache. Self-contained — no consumers yet.
Files:
Create: /Users/buoy/Development/gitrepo/PPT/src/views/Editor/EnglishSpeaking/composables/useAudioPlayer.ts
[ ] Step 1: Create the composable
Create src/views/Editor/EnglishSpeaking/composables/useAudioPlayer.ts:
import { ref, onUnmounted, type Ref } from 'vue'
import { synthesize } from '../services/speechService'
export type PlaySource =
| { kind: 'tts'; text: string }
| { kind: 'blob'; blob: Blob }
export interface AudioPlayer {
/** Id of the message currently playing (audio element fired `playing`). */
playingId: Readonly<Ref<string | null>>
/** Id of the message whose synthesis or play() is in flight. */
loadingId: Readonly<Ref<string | null>>
/** Id of the message whose last playback attempt failed. Sticky until next play(). */
errorId: Readonly<Ref<string | null>>
play(id: string, source: PlaySource): Promise<void>
stop(): void
}
/**
* Single audio playback owner for the dialogue view.
*
* Structural guarantees:
* - At most one HTMLAudioElement / one in-flight synthesis at a time.
* Every play() begins by aborting & pausing the prior session.
* - Cached MP3 blobs (one per AI messageId) live in-memory; URLs are
* revoked on view unmount.
* - Errors collapse to a single per-id state. The view renders a retry
* affordance; clicking the same play button re-enters play().
*/
export function useAudioPlayer(): AudioPlayer {
const playingId = ref<string | null>(null)
const loadingId = ref<string | null>(null)
const errorId = ref<string | null>(null)
// Closure-private state. Not reactive on purpose.
let currentAudio: HTMLAudioElement | null = null
let synthAbort: AbortController | null = null
const ttsCache = new Map<string, Blob>()
const cachedUrls: string[] = []
function clearCurrentAudio() {
if (currentAudio) {
currentAudio.onplaying = null
currentAudio.onended = null
currentAudio.onerror = null
try { currentAudio.pause() } catch { /* ignore */ }
currentAudio = null
}
}
function failPlayback(id: string, reason: unknown) {
if (loadingId.value === id) loadingId.value = null
if (playingId.value === id) playingId.value = null
errorId.value = id
console.warn('[audio-player] playback failed:', id, reason)
}
async function play(id: string, source: PlaySource): Promise<void> {
// A fresh attempt — drop stale error.
errorId.value = null
// Abort any in-flight synthesis and pause any current audio.
stop()
loadingId.value = id
try {
let blob: Blob
if (source.kind === 'blob') {
blob = source.blob
}
else {
const cached = ttsCache.get(id)
if (cached) {
blob = cached
}
else {
synthAbort = new AbortController()
blob = await synthesize(source.text, synthAbort.signal)
synthAbort = null
// We may have been interrupted while awaiting (loadingId changed).
if (loadingId.value !== id) return
ttsCache.set(id, blob)
}
}
const url = URL.createObjectURL(blob)
cachedUrls.push(url)
const audio = new Audio(url)
currentAudio = audio
audio.onplaying = () => {
if (loadingId.value === id) loadingId.value = null
playingId.value = id
}
audio.onended = () => {
if (currentAudio === audio) currentAudio = null
if (playingId.value === id) playingId.value = null
}
audio.onerror = () => {
if (currentAudio === audio) currentAudio = null
failPlayback(id, audio.error)
}
try {
await audio.play()
}
catch (err) {
// Most often: autoplay policy blocked the call.
if (currentAudio === audio) currentAudio = null
failPlayback(id, err)
}
}
catch (err) {
// Synthesis path errors. AbortError fires when stop() was called
// mid-synthesis — that is a normal interrupt, not a failure.
synthAbort = null
if (err instanceof Error && err.name === 'AbortError') {
if (loadingId.value === id) loadingId.value = null
return
}
failPlayback(id, err)
}
}
function stop(): void {
if (synthAbort) {
synthAbort.abort()
synthAbort = null
}
clearCurrentAudio()
loadingId.value = null
playingId.value = null
// Note: errorId is intentionally left alone here. It is only
// cleared at the start of a new play() attempt.
}
onUnmounted(() => {
stop()
for (const url of cachedUrls) {
try { URL.revokeObjectURL(url) } catch { /* ignore */ }
}
cachedUrls.length = 0
ttsCache.clear()
})
return {
playingId,
loadingId,
errorId,
play,
stop,
}
}
Run:
cd /Users/buoy/Development/gitrepo/PPT
npm run type-check
Expected: exits 0.
cd /Users/buoy/Development/gitrepo/PPT
git add src/views/Editor/EnglishSpeaking/composables/useAudioPlayer.ts
git commit -m "$(cat <<'EOF'
feat: add useAudioPlayer composable
Sole audio playback owner with structural single-channel
guarantee, per-id error surface, and synthesis result cache.
Calls speechService.synthesize for kind:'tts' sources and
plays kind:'blob' sources directly.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
EOF
)"
Strip the engine of all audio concerns and add the new helper. Keep build green by removing engine.cancelTTS() from the one place the view calls it (handleRestart).
Files:
/Users/buoy/Development/gitrepo/PPT/src/views/Editor/EnglishSpeaking/composables/useDialogueEngine.tsModify: /Users/buoy/Development/gitrepo/PPT/src/views/Editor/EnglishSpeaking/preview/DialogueChatView.vue
[ ] Step 1: Remove ttsUtterance declaration
In useDialogueEngine.ts, delete this line (currently around line 16):
let ttsUtterance: SpeechSynthesisUtterance | null = null
speakTTS call inside generateGreetingIn the same file, inside generateGreeting(), delete the line:
speakTTS(aiMessage)
(Currently around line 56, between aiMsg.status = 'done' and the catch.)
speakTTS call inside sendStudentMessageInside sendStudentMessage()'s 'done' event branch, delete:
speakTTS(aiMsg.content)
(Currently around line 124.)
speakTTS call inside regenerateAiMessageInside regenerateAiMessage()'s 'done' event branch, delete:
speakTTS(aiMsg.content)
(Currently around line 201.)
speakTTS call inside beginStudentStreamInside beginStudentStream()'s ws.onmessage done branch, delete:
speakTTS(aiMsg.content)
(Currently around line 369.)
speakTTS and cancelTTS function definitionsDelete the entire ==================== TTS ==================== section, currently around lines 242–259:
// ==================== TTS ====================
function speakTTS(text: string) {
if (!text || typeof speechSynthesis === 'undefined') return
cancelTTS()
ttsUtterance = new SpeechSynthesisUtterance(text)
ttsUtterance.lang = 'en-US'
ttsUtterance.rate = 0.9
speechSynthesis.speak(ttsUtterance)
}
function cancelTTS() {
if (typeof speechSynthesis !== 'undefined') {
speechSynthesis.cancel()
}
ttsUtterance = null
}
cancelTTS() from onUnmountedInside the onUnmounted block (currently around line 426–431), delete the cancelTTS() line:
onUnmounted(() => {
abort()
greetingAbortController?.abort()
cancelTTS() // ← remove this line
stopCountdown()
})
The block becomes:
onUnmounted(() => {
abort()
greetingAbortController?.abort()
stopCountdown()
})
attachStudentBlob helperInsert this function after streamFallback and before the // ==================== Cleanup ==================== marker:
/**
* Attach the recorded audio blob to a student message that was
* pushed by `beginStudentStream` (which doesn't know the final
* blob until the user clicks 完成). Lets click-replay work in
* the WebSocket path the same way it does for HTTP.
*/
function attachStudentBlob(messageId: string, blob: Blob) {
const msg = messages.value.find(m => m.id === messageId)
if (msg && msg.role === 'student') msg.audioBlob = blob
}
In the return { ... } at the bottom of useDialogueEngine, replace cancelTTS with attachStudentBlob. The block currently ends:
streamFallback,
retryMessage,
regenerateAiMessage,
getReport,
completeSession,
abort,
cancelTTS,
}
Change to:
streamFallback,
retryMessage,
regenerateAiMessage,
getReport,
completeSession,
abort,
attachStudentBlob,
}
engine.cancelTTS() from the view's handleRestartIn DialogueChatView.vue, inside handleRestart (currently around line 801–806), delete the engine.cancelTTS() line:
function handleRestart() {
showExitConfirm.value = false
engine.abort()
engine.cancelTTS() // ← remove this line
emit('restart')
}
The function becomes:
function handleRestart() {
showExitConfirm.value = false
engine.abort()
emit('restart')
}
(player.stop() is wired in here in Task 4.)
Run:
cd /Users/buoy/Development/gitrepo/PPT
npm run type-check
Expected: exits 0. There is one transient regression in this commit — auto-play of AI replies stops working. That is restored in Task 5 (auto-play watcher). The build itself is green.
cd /Users/buoy/Development/gitrepo/PPT
git add src/views/Editor/EnglishSpeaking/composables/useDialogueEngine.ts src/views/Editor/EnglishSpeaking/preview/DialogueChatView.vue
git commit -m "$(cat <<'EOF'
refactor: remove TTS ownership from dialogue engine
Drop speakTTS / cancelTTS / ttsUtterance and all four call
sites — the engine no longer touches speechSynthesis. Add
attachStudentBlob so the WebSocket path can attach the final
recorded blob to its student message. Drop the now-unused
engine.cancelTTS() call from DialogueChatView.handleRestart.
Auto-play is restored in a follow-up commit that wires the
new view-level audio player.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
EOF
)"
Replace all view-local audio state (currentAudio, currentAudioUrl, stopCurrentPlayback, playingMessageId) with the new player. Rewrite togglePlay against player.play / player.stop. Fix the WS path's missing audioBlob. Also wire player.stop() into handleStartRecording and handleRestart. Auto-play is still NOT wired — that comes in Task 5.
Files:
Modify: /Users/buoy/Development/gitrepo/PPT/src/views/Editor/EnglishSpeaking/preview/DialogueChatView.vue
[ ] Step 1: Import the player composable
In the <script setup> imports near the top (currently around line 425–432), add:
import { useAudioPlayer } from '../composables/useAudioPlayer'
(Place it next to the other composable imports — after useAudioRecorder.)
In the Composables block (currently around line 484–488), add player:
const engine = useDialogueEngine()
const recorder = useAudioRecorder()
const player = useAudioPlayer()
Delete these declarations from the Local UI State block (currently around line 495 and lines 512–513):
const playingMessageId = ref<string | null>(null)
let currentAudio: HTMLAudioElement | null = null
let currentAudioUrl: string | null = null
stopCurrentPlaybackDelete the entire function (currently around lines 716–727):
function stopCurrentPlayback() {
if (currentAudio) {
currentAudio.pause()
currentAudio = null
}
if (currentAudioUrl) {
URL.revokeObjectURL(currentAudioUrl)
currentAudioUrl = null
}
if (typeof speechSynthesis !== 'undefined') speechSynthesis.cancel()
playingMessageId.value = null
}
togglePlayReplace the entire togglePlay(id) function (currently around lines 729–770) with:
function togglePlay(id: string) {
// Same id is currently playing or loading → stop.
if (player.playingId.value === id || player.loadingId.value === id) {
player.stop()
return
}
const msg = engine.messages.value.find(m => m.id === id)
if (!msg) return
if (msg.role === 'student' && msg.audioBlob) {
player.play(id, { kind: 'blob', blob: msg.audioBlob })
}
else if (msg.role === 'ai' && msg.content) {
player.play(id, { kind: 'tts', text: msg.content })
}
}
Update handleStartRecording (currently around lines 608–624). Add player.stop() immediately after the early-return guard:
async function handleStartRecording() {
if (!engine.canRecord.value || recorder.isRecording.value) return
player.stop() // ← new
try {
await recorder.startRecording()
// ...rest unchanged
Update handleRestart (currently — after Task 3 — three lines):
function handleRestart() {
showExitConfirm.value = false
engine.abort()
player.stop() // ← new
emit('restart')
}
Update handleFinishRecording (currently around lines 639–656). The if (ctl) branch needs to attach the blob before calling ctl.finish():
async function handleFinishRecording() {
if (!recorder.isRecording.value) return
const ctl = streamCtl
streamCtl = null
try {
const blob = await recorder.stopRecording()
recorder.onChunk.value = null
if (ctl) {
engine.attachStudentBlob(ctl.studentMsgId, blob) // ← new
ctl.finish()
}
else {
await engine.sendStudentMessage(blob)
}
}
catch (err) {
console.error('Recording/send failed:', err)
}
}
stopCurrentPlayback() call from onUnmountedThe composable's own onUnmounted already calls stop() and revokes URLs. In DialogueChatView.vue's onUnmounted (currently around lines 927–931), remove the stopCurrentPlayback() line:
onUnmounted(() => {
if (idleHintTimer) clearTimeout(idleHintTimer)
if (badgeTimer) clearTimeout(badgeTimer)
stopCurrentPlayback() // ← remove
})
The block becomes:
onUnmounted(() => {
if (idleHintTimer) clearTimeout(idleHintTimer)
if (badgeTimer) clearTimeout(badgeTimer)
})
playingMessageIdThe template currently uses playingMessageId in two places (around lines 52 and 108) inside the play-button SVGs:
<svg v-if="playingMessageId !== message.id" ...>
These conditions will be replaced fully in Task 6. For now, just keep the build green — replace each playingMessageId !== message.id with player.playingId.value !== message.id:
AI message play button (line ~52):
<button class="play-btn play-ai" @click="togglePlay(message.id)">
<svg v-if="player.playingId.value !== message.id" width="12" height="12" viewBox="0 0 24 24" fill="currentColor">
<polygon points="5 3 19 12 5 21 5 3" />
</svg>
<svg v-else width="12" height="12" viewBox="0 0 24 24" fill="currentColor">
<rect x="6" y="4" width="4" height="16" /><rect x="14" y="4" width="4" height="16" />
</svg>
</button>
Student message play button (line ~108):
<button class="play-btn play-student" @click="togglePlay(message.id)">
<svg v-if="player.playingId.value !== message.id" width="12" height="12" viewBox="0 0 24 24" fill="currentColor">
<polygon points="5 3 19 12 5 21 5 3" />
</svg>
<svg v-else width="12" height="12" viewBox="0 0 24 24" fill="currentColor">
<rect x="6" y="4" width="4" height="16" /><rect x="14" y="4" width="4" height="16" />
</svg>
</button>
(Loading and error states are added in Task 6; this step is a holding edit so the file compiles.)
Run:
cd /Users/buoy/Development/gitrepo/PPT
npm run type-check
Expected: exits 0.
Start dev server in another terminal: npm run dev. Open the dialogue view. Verify:
*.tts.speech.microsoft.com) and plays.Note: AI auto-play after model done does NOT yet fire — that is Task 5.
cd /Users/buoy/Development/gitrepo/PPT
git add src/views/Editor/EnglishSpeaking/preview/DialogueChatView.vue
git commit -m "$(cat <<'EOF'
refactor: route DialogueChatView click-playback through useAudioPlayer
Replace view-local currentAudio / stopCurrentPlayback /
playingMessageId with the new player composable. togglePlay
now dispatches student blobs and AI text to player.play, and
handleStartRecording / handleRestart call player.stop so
recording always interrupts audio. WebSocket path now attaches
the recorded blob to its student message (engine.attachStudentBlob)
so click-replay works in WS mode.
Auto-play of new AI replies is restored in the next commit.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
EOF
)"
Restore auto-play of AI replies after done. View watches engine.messages for transitions to (role: 'ai', status: 'done') and fires player.play(id, { kind:'tts', text }), deduplicated via a per-component Set<string>.
Files:
Modify: /Users/buoy/Development/gitrepo/PPT/src/views/Editor/EnglishSpeaking/preview/DialogueChatView.vue
[ ] Step 1: Add the dedup set + watcher
In the Watchers block (after the existing watcher that scrolls to bottom — currently the first watch(...) around line 844), insert:
// 自动播放:AI 消息流式 done 后,合成并播一次。
// 用 Set 去重防止 watcher 因为不相关重渲染重复触发。
const autoPlayedIds = new Set<string>()
watch(
() => engine.messages.value.map(m => `${m.id}:${m.status}`).join('|'),
() => {
for (const m of engine.messages.value) {
if (
m.role === 'ai' &&
m.status === 'done' &&
m.content &&
!autoPlayedIds.has(m.id)
) {
autoPlayedIds.add(m.id)
player.play(m.id, { kind: 'tts', text: m.content })
}
}
},
)
(Place it next to the other watch(...) blocks. Order does not matter functionally.)
Run:
cd /Users/buoy/Development/gitrepo/PPT
npm run type-check
Expected: exits 0.
npm run dev → open dialogue view. Verify:
playingId === id).Re-generating an AI error message: the new id auto-plays again (different id, dedup set doesn't block).
[ ] Step 4: Commit
cd /Users/buoy/Development/gitrepo/PPT
git add src/views/Editor/EnglishSpeaking/preview/DialogueChatView.vue
git commit -m "$(cat <<'EOF'
feat: auto-play AI replies via useAudioPlayer
Watch engine.messages for AI messages reaching status='done'
and dispatch them to player.play once each. A per-component
Set<string> deduplicates so unrelated reactive churn does not
re-trigger playback.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
EOF
)"
Render idle / loading / playing / error on every voice-bar play button, swap the duration label for "点击重试" when in error, and add CSS for the loading spinner + error variant.
Files:
Modify: /Users/buoy/Development/gitrepo/PPT/src/views/Editor/EnglishSpeaking/preview/DialogueChatView.vue
[ ] Step 1: Replace the AI message play button
Locate the AI voice-bar play button (currently around lines 51–58). Replace the entire <button> element with:
<button
class="play-btn play-ai"
:class="{ 'play-btn-error': player.errorId.value === message.id }"
@click="togglePlay(message.id)"
>
<svg
v-if="player.loadingId.value === message.id"
class="play-spinner"
width="12" height="12" viewBox="0 0 24 24" fill="none" stroke="currentColor"
stroke-width="2" stroke-linecap="round"
>
<path d="M21 12a9 9 0 1 1-6.219-8.56" />
</svg>
<svg
v-else-if="player.playingId.value === message.id"
width="12" height="12" viewBox="0 0 24 24" fill="currentColor"
>
<rect x="6" y="4" width="4" height="16" /><rect x="14" y="4" width="4" height="16" />
</svg>
<svg
v-else-if="player.errorId.value === message.id"
width="12" height="12" viewBox="0 0 24 24" fill="none" stroke="currentColor"
stroke-width="2" stroke-linecap="round" stroke-linejoin="round"
>
<path d="M12 9v4" />
<path d="M12 17h.01" />
<path d="M10.29 3.86 1.82 18a2 2 0 0 0 1.71 3h16.94a2 2 0 0 0 1.71-3L13.71 3.86a2 2 0 0 0-3.42 0z" />
</svg>
<svg
v-else
width="12" height="12" viewBox="0 0 24 24" fill="currentColor"
>
<polygon points="5 3 19 12 5 21 5 3" />
</svg>
</button>
Immediately after the <div class="wave-bar-group">…</div> block in the AI voice-bar (currently around line 67), replace:
<span class="voice-duration voice-duration-ai">0:04</span>
with:
<span
v-if="player.errorId.value === message.id"
class="play-error-hint"
>点击重试</span>
<span
v-else
class="voice-duration voice-duration-ai"
>0:04</span>
Locate the student voice-bar play button (currently around lines 107–114). Replace the entire <button> element with:
<button
class="play-btn play-student"
:class="{ 'play-btn-error': player.errorId.value === message.id }"
@click="togglePlay(message.id)"
>
<svg
v-if="player.loadingId.value === message.id"
class="play-spinner"
width="12" height="12" viewBox="0 0 24 24" fill="none" stroke="currentColor"
stroke-width="2" stroke-linecap="round"
>
<path d="M21 12a9 9 0 1 1-6.219-8.56" />
</svg>
<svg
v-else-if="player.playingId.value === message.id"
width="12" height="12" viewBox="0 0 24 24" fill="currentColor"
>
<rect x="6" y="4" width="4" height="16" /><rect x="14" y="4" width="4" height="16" />
</svg>
<svg
v-else-if="player.errorId.value === message.id"
width="12" height="12" viewBox="0 0 24 24" fill="none" stroke="currentColor"
stroke-width="2" stroke-linecap="round" stroke-linejoin="round"
>
<path d="M12 9v4" />
<path d="M12 17h.01" />
<path d="M10.29 3.86 1.82 18a2 2 0 0 0 1.71 3h16.94a2 2 0 0 0 1.71-3L13.71 3.86a2 2 0 0 0-3.42 0z" />
</svg>
<svg
v-else
width="12" height="12" viewBox="0 0 24 24" fill="currentColor"
>
<polygon points="5 3 19 12 5 21 5 3" />
</svg>
</button>
Find the student voice-bar duration <span> (currently around line 98):
<span class="voice-duration voice-duration-student">0:04</span>
Replace with:
<span
v-if="player.errorId.value === message.id"
class="play-error-hint play-error-hint-student"
>点击重试</span>
<span
v-else
class="voice-duration voice-duration-student"
>0:04</span>
In the <style lang="scss" scoped> block, locate the .play-student rule (currently around line 1093). Append the new styles immediately after .play-student { ... } (and before .wave-bar-group):
.play-btn-error {
background: #fef2f2 !important;
color: #dc2626 !important;
border: 1px solid #fecaca;
&:hover { background: #fee2e2 !important; }
}
.play-spinner {
animation: spin 1s linear infinite;
}
.play-error-hint {
font-size: 10px;
color: #dc2626;
font-weight: 500;
flex-shrink: 0;
white-space: nowrap;
}
.play-error-hint-student {
color: #fff;
}
(@keyframes spin already exists later in this file — no need to redeclare.)
Run:
cd /Users/buoy/Development/gitrepo/PPT
npm run type-check
npm run build
Both commands: exits 0.
npm run dev → open dialogue view.
Click an active playing message: button toggles to ▶ instantly.
[ ] Step 8: Manual verification — error paths
Set a bogus VITE_AZURE_SPEECH_KEY in .env.local and restart dev server.
Disconnect network mid-conversation:
Reconnect → click button → plays.
[ ] Step 9: Manual verification — single-channel + interrupt
AI auto-playing → click another AI message's button → first stops, second plays.
AI auto-playing → click a student message's button → first stops, student plays.
Student replaying → AI auto-play arrives → student stops, AI plays.
AI auto-playing → click 开始录音 → audio stops within 100ms.
Synthesis in progress (slow network — throttle to "Slow 3G") → click 开始录音 → spinner clears immediately, no audio plays.
[ ] Step 10: Manual verification — cache + lifecycle
Open DevTools Network panel filtered to tts.speech.microsoft.com.
Navigate out of the dialogue view → back in → no zombie audio. No "URL revoked" console warnings.
[ ] Step 11: Commit
cd /Users/buoy/Development/gitrepo/PPT
git add src/views/Editor/EnglishSpeaking/preview/DialogueChatView.vue
git commit -m "$(cat <<'EOF'
feat: render player state on voice-bar play buttons
Each play button derives idle / loading / playing / error from
the player refs. Errors swap "0:04" for "点击重试"; clicking
the warning button retries via the same togglePlay path. AI and
student variants share the same logic with the existing color
schemes preserved.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
EOF
)"
Spec coverage:
Placeholder scan: No "TBD", no "implement later", no "similar to Task N", no unspecified error handlers — every step has full code.
Type consistency: play(id, source), stop(), playingId/loadingId/errorId, attachStudentBlob(messageId, blob), PlaySource discriminated union — used identically across the composable definition (T2), engine signature (T3), and view consumers (T4, T5, T6). Method calls in the view (player.play / player.stop / player.playingId.value) match the composable's exported shape. The engine's new helper signature matches its single call site. No mismatches found.
Plan complete and saved to docs/superpowers/plans/2026-04-26-azure-tts-audio-player.md. Two execution options:
1. Subagent-Driven (recommended) — I dispatch a fresh subagent per task, review between tasks, fast iteration
2. Inline Execution — Execute tasks in this session using executing-plans, batch execution with checkpoints
Which approach?