Commit Β·
e7db016
1
Parent(s): 4f6f969
feat: Added voice journal recording with Cohere ASR integration
Browse files- COHERE_ASR_SETUP.md +127 -0
- README.md +14 -1
- app.py +85 -6
- app/data/sessions_store.json +1 -1
- app/schemas/journal_schema.json +32 -0
- app/services/asr.py +269 -0
- app/services/journal.py +48 -6
- requirements.txt +5 -0
- test_asr.py +249 -0
COHERE_ASR_SETUP.md
ADDED
|
@@ -0,0 +1,127 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Cohere Transcribe ASR β Setup Guide
|
| 2 |
+
|
| 3 |
+
## Overview
|
| 4 |
+
|
| 5 |
+
The voice journal pipeline uses **CohereLabs/cohere-transcribe-03-2026**, a
|
| 6 |
+
2 B-parameter conformer encoder + lightweight transformer decoder trained
|
| 7 |
+
from scratch for ASR. It is gated on the Hugging Face Hub (you must
|
| 8 |
+
accept the model terms once with your account) and supports 14
|
| 9 |
+
languages:
|
| 10 |
+
|
| 11 |
+
- **European:** English, French, German, Italian, Spanish, Portuguese,
|
| 12 |
+
Greek, Dutch, Polish
|
| 13 |
+
- **APAC:** Chinese (Mandarin), Japanese, Korean, Vietnamese
|
| 14 |
+
- **MENA:** Arabic
|
| 15 |
+
|
| 16 |
+
The model is Apache 2.0 licensed and is integrated into the journal
|
| 17 |
+
pipeline so players can speak during a game instead of typing.
|
| 18 |
+
|
| 19 |
+
## Why this configuration
|
| 20 |
+
|
| 21 |
+
1. **Sponsor visibility** β Cohere Labs is a hackathon sponsor.
|
| 22 |
+
2. **State-of-the-art accuracy** β 5.42 mean WER on the Open ASR
|
| 23 |
+
Leaderboard (5.xβ10.x WER across real-world domains) and 1.25 WER on
|
| 24 |
+
LibriSpeech clean.
|
| 25 |
+
3. **Production runtime** β supports π€ Transformers (offline),
|
| 26 |
+
vLLM, mlx-audio, Rust, and a WebGPU browser demo.
|
| 27 |
+
4. **Lazy loading** β the model is downloaded on first use, never at
|
| 28 |
+
app startup, so demo boot is unaffected.
|
| 29 |
+
|
| 30 |
+
## Installation
|
| 31 |
+
|
| 32 |
+
### 1. Accept the model terms
|
| 33 |
+
|
| 34 |
+
Visit <https://huggingface.co/CohereLabs/cohere-transcribe-03-2026>,
|
| 35 |
+
click **Agree and access repository** with the account you plan to
|
| 36 |
+
authenticate as.
|
| 37 |
+
|
| 38 |
+
### 2. Install dependencies
|
| 39 |
+
|
| 40 |
+
```bash
|
| 41 |
+
pip install 'transformers>=5.4.0' torch huggingface_hub \
|
| 42 |
+
soundfile librosa sentencepiece protobuf
|
| 43 |
+
```
|
| 44 |
+
|
| 45 |
+
(These are added to `requirements.txt` for the demo; `transformers` is
|
| 46 |
+
already present.)
|
| 47 |
+
|
| 48 |
+
### 3. Provide an HF token
|
| 49 |
+
|
| 50 |
+
Set a token in the environment so the gated model can be downloaded:
|
| 51 |
+
|
| 52 |
+
```bash
|
| 53 |
+
export HF_TOKEN=hf_xxx... # Linux/macOS
|
| 54 |
+
$env:HF_TOKEN="hf_xxx..." # PowerShell
|
| 55 |
+
```
|
| 56 |
+
|
| 57 |
+
On Hugging Face Spaces, create a `HF_TOKEN` secret (same name as
|
| 58 |
+
`huggingface` in `modal_serve.py`/`modal_train.py`).
|
| 59 |
+
|
| 60 |
+
## How the pipeline uses it
|
| 61 |
+
|
| 62 |
+
```text
|
| 63 |
+
Gradio microphone / upload
|
| 64 |
+
β
|
| 65 |
+
app.py: record_journal(audio_path, language)
|
| 66 |
+
β
|
| 67 |
+
app/services/asr.py: transcribe(audio_path, language)
|
| 68 |
+
β
|
| 69 |
+
CohereAsrForConditionalGeneration β CohereLabs/cohere-transcribe-03-2026
|
| 70 |
+
β
|
| 71 |
+
transcript
|
| 72 |
+
β
|
| 73 |
+
app/services/journal.py: create_journal_entry(...)
|
| 74 |
+
β
|
| 75 |
+
app/logs/journals.jsonl
|
| 76 |
+
```
|
| 77 |
+
|
| 78 |
+
Each journal entry now carries:
|
| 79 |
+
|
| 80 |
+
- `transcript_source` β `"typed" | "asr" | "hybrid"`
|
| 81 |
+
- `audio_ref` β path of the recorded audio clip
|
| 82 |
+
- `asr` β `{ model, language, status, error }`
|
| 83 |
+
|
| 84 |
+
The `journal_recorded` event log also includes `transcript_source`,
|
| 85 |
+
`asr_status`, and `asr_model` for full traceability.
|
| 86 |
+
|
| 87 |
+
## Skipping the model in tests
|
| 88 |
+
|
| 89 |
+
To run the demo without downloading the model, set either of:
|
| 90 |
+
|
| 91 |
+
```bash
|
| 92 |
+
CITYQUEST_SKIP_MODEL=1
|
| 93 |
+
CITYQUEST_FAST_TEST=1
|
| 94 |
+
```
|
| 95 |
+
|
| 96 |
+
When set, `app.services.asr.transcribe()` returns
|
| 97 |
+
`status="skipped"` with an empty transcript. The journal pipeline
|
| 98 |
+
silently falls back to typed input.
|
| 99 |
+
|
| 100 |
+
## Verification
|
| 101 |
+
|
| 102 |
+
```bash
|
| 103 |
+
$env:CITYQUEST_FAST_TEST="1"
|
| 104 |
+
.\.venv\Scripts\python.exe test_asr.py
|
| 105 |
+
.\.venv\Scripts\python.exe test_end_to_end.py
|
| 106 |
+
```
|
| 107 |
+
|
| 108 |
+
Expected: 30/30 ASR tests + 86/86 end-to-end tests pass in skip-mode.
|
| 109 |
+
|
| 110 |
+
## Limitations (per the model card)
|
| 111 |
+
|
| 112 |
+
1. **Single language per call** β pick the right language code; the
|
| 113 |
+
model does not auto-detect or handle code-switching well.
|
| 114 |
+
2. **No diarization or timestamps** β only plain text is returned.
|
| 115 |
+
3. **Eager on silence** β prepend a VAD/silence gate if the recording
|
| 116 |
+
has noisy backgrounds; otherwise the model may hallucinate.
|
| 117 |
+
|
| 118 |
+
## File map
|
| 119 |
+
|
| 120 |
+
| File | Purpose |
|
| 121 |
+
| --- | --- |
|
| 122 |
+
| `app/services/asr.py` | Lazy-loaded Cohere Transcribe wrapper. |
|
| 123 |
+
| `app/services/journal.py` | `transcribe_journal()` and `create_journal_entry()` now accept ASR metadata. |
|
| 124 |
+
| `app/schemas/journal_schema.json` | Optional `transcript_source`, `audio_ref`, `asr` fields. |
|
| 125 |
+
| `app.py` | Gradio audio component + `record_journal()` voice path. |
|
| 126 |
+
| `test_asr.py` | Skip-mode tests for the ASR pipeline. |
|
| 127 |
+
| `requirements.txt` | Optional ASR runtime deps. |
|
README.md
CHANGED
|
@@ -28,7 +28,20 @@ An AI-powered platform that generates complete, playable real-world games β ru
|
|
| 28 |
1. Select a game type, city, and preferences
|
| 29 |
2. AI generates rules, tasks, and hints
|
| 30 |
3. Play in the real world, track progress in the app
|
| 31 |
-
4.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 32 |
|
| 33 |
## Team
|
| 34 |
|
|
|
|
| 28 |
1. Select a game type, city, and preferences
|
| 29 |
2. AI generates rules, tasks, and hints
|
| 30 |
3. Play in the real world, track progress in the app
|
| 31 |
+
4. Record voice journals during play (auto-transcribed with [Cohere Transcribe](https://huggingface.co/CohereLabs/cohere-transcribe-03-2026))
|
| 32 |
+
5. Receive an AI-generated story summary of the outcome
|
| 33 |
+
|
| 34 |
+
## AI Models
|
| 35 |
+
|
| 36 |
+
| Stage | Model | Notes |
|
| 37 |
+
| --- | --- | --- |
|
| 38 |
+
| Game generation | `nvidia/NVIDIA-Nemotron-3-Nano-4B-GGUF` via llama.cpp | Retrieval-grounded, validation + repair enforced |
|
| 39 |
+
| Voice journal ASR | `CohereLabs/cohere-transcribe-03-2026` via π€ Transformers | 14 languages, lazy-loaded; typed-input fallback |
|
| 40 |
+
| Recap (optional) | `openbmb/MiniCPM5-1B-GGUF` | Currently using deterministic template recap for reliability |
|
| 41 |
+
| Poster (optional) | `black-forest-labs/FLUX.1-schnell` | Skipped under `CITYQUEST_SKIP_MODEL` |
|
| 42 |
+
|
| 43 |
+
Set `CITYQUEST_FAST_TEST=1` to run the entire pipeline without
|
| 44 |
+
downloading any model weights.
|
| 45 |
|
| 46 |
## Team
|
| 47 |
|
app.py
CHANGED
|
@@ -215,22 +215,77 @@ def use_hint(session_id: str, task_id: str, team_id: str = "team-a"):
|
|
| 215 |
|
| 216 |
def record_journal(
|
| 217 |
session_id: str,
|
| 218 |
-
transcript: str,
|
| 219 |
task_id: str = "",
|
| 220 |
location_note: str = "",
|
| 221 |
team_id: str = "team-a",
|
|
|
|
|
|
|
| 222 |
):
|
| 223 |
-
"""Record a
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 224 |
if session_id not in SESSION_STORE:
|
| 225 |
return "[!] Unknown session"
|
| 226 |
|
| 227 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 228 |
entry = create_journal_entry(
|
| 229 |
transcript=transcript,
|
| 230 |
session_id=session_id,
|
| 231 |
team_id=team_id,
|
| 232 |
task_id=task_id or None,
|
| 233 |
location_note=location_note,
|
|
|
|
|
|
|
|
|
|
| 234 |
)
|
| 235 |
|
| 236 |
# Summarize
|
|
@@ -248,16 +303,28 @@ def record_journal(
|
|
| 248 |
"mood": entry["mood"],
|
| 249 |
"story_value": summary["story_value"],
|
| 250 |
"summary": summary["moment_summary"],
|
|
|
|
|
|
|
|
|
|
| 251 |
}, team_id=team_id)
|
| 252 |
SESSION_STORE[session_id]["events"].append(ev)
|
| 253 |
SESSION_STORE[session_id]["journals"].append(entry)
|
| 254 |
|
| 255 |
# Build display text
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 256 |
display = (
|
| 257 |
-
f"ποΈ **Journal recorded!**\n"
|
| 258 |
f"- Mood: *{entry['mood']}*\n"
|
| 259 |
f"- Story value: **{summary['story_value']}**\n"
|
| 260 |
f"- Tags: {', '.join(summary['tags'])}\n"
|
|
|
|
| 261 |
f"- Summary: {summary['moment_summary']}"
|
| 262 |
)
|
| 263 |
return display
|
|
@@ -1008,13 +1075,24 @@ with gr.Blocks(title="CityQuest-AI") as demo:
|
|
| 1008 |
task_feedback = gr.Markdown(value="")
|
| 1009 |
|
| 1010 |
gr.HTML('<div class="section-label" style="margin-top:20px;">Voice Journal</div>')
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1011 |
journal_transcript = gr.Textbox(
|
| 1012 |
-
label="
|
|
|
|
| 1013 |
placeholder="We just found the mural near the canal β it was incredible!",
|
| 1014 |
)
|
| 1015 |
with gr.Row():
|
| 1016 |
journal_task_id = gr.Textbox(label="Task ID (opt.)", placeholder="t1")
|
| 1017 |
journal_location = gr.Textbox(label="Location note", placeholder="Rue de Rivoli")
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1018 |
journal_btn = gr.Button("ποΈ Record Journal", variant="secondary")
|
| 1019 |
journal_output = gr.Markdown(value="")
|
| 1020 |
|
|
@@ -1226,7 +1304,8 @@ with gr.Blocks(title="CityQuest-AI") as demo:
|
|
| 1226 |
|
| 1227 |
journal_btn.click(
|
| 1228 |
fn=record_journal,
|
| 1229 |
-
inputs=[current_session, journal_transcript, journal_task_id, journal_location,
|
|
|
|
| 1230 |
outputs=[journal_output],
|
| 1231 |
)
|
| 1232 |
photo_btn.click(
|
|
|
|
| 215 |
|
| 216 |
def record_journal(
|
| 217 |
session_id: str,
|
| 218 |
+
transcript: str = "",
|
| 219 |
task_id: str = "",
|
| 220 |
location_note: str = "",
|
| 221 |
team_id: str = "team-a",
|
| 222 |
+
audio_path: str = "",
|
| 223 |
+
language: str = "en",
|
| 224 |
):
|
| 225 |
+
"""Record a journal entry, summarize it, and return the result.
|
| 226 |
+
|
| 227 |
+
Two input paths are supported:
|
| 228 |
+
|
| 229 |
+
* **Voice path** β pass ``audio_path`` (e.g. a path returned by
|
| 230 |
+
``gr.Audio(type="filepath")``). The audio is transcribed with the
|
| 231 |
+
Cohere ASR service (``CohereLabs/cohere-transcribe-03-2026``) and
|
| 232 |
+
the transcript is stored with ASR metadata.
|
| 233 |
+
* **Typed path** β pass ``transcript`` directly. The user can also
|
| 234 |
+
edit an ASR transcript in the UI before submitting; if both
|
| 235 |
+
``audio_path`` and ``transcript`` are present, the typed
|
| 236 |
+
transcript wins and is treated as a manual correction
|
| 237 |
+
(``transcript_source == "hybrid"``).
|
| 238 |
+
"""
|
| 239 |
if session_id not in SESSION_STORE:
|
| 240 |
return "[!] Unknown session"
|
| 241 |
|
| 242 |
+
asr_metadata: dict | None = None
|
| 243 |
+
audio_ref: str | None = None
|
| 244 |
+
transcript_source = "typed"
|
| 245 |
+
|
| 246 |
+
# ββ 1. Voice path β transcribe audio if provided βββββββββββββββββββ
|
| 247 |
+
if audio_path:
|
| 248 |
+
audio_ref = audio_path
|
| 249 |
+
try:
|
| 250 |
+
from app.services.asr import transcribe as _asr_transcribe
|
| 251 |
+
|
| 252 |
+
asr_result = _asr_transcribe(audio_path, language=language)
|
| 253 |
+
asr_metadata = {
|
| 254 |
+
"model": asr_result.get("model"),
|
| 255 |
+
"language": asr_result.get("language"),
|
| 256 |
+
"status": asr_result.get("status"),
|
| 257 |
+
"error": asr_result.get("error"),
|
| 258 |
+
}
|
| 259 |
+
asr_text = (asr_result.get("transcript") or "").strip()
|
| 260 |
+
if asr_text and not transcript.strip():
|
| 261 |
+
transcript = asr_text
|
| 262 |
+
transcript_source = "asr"
|
| 263 |
+
elif asr_text and transcript.strip() and asr_text != transcript.strip():
|
| 264 |
+
transcript_source = "hybrid"
|
| 265 |
+
except Exception as exc:
|
| 266 |
+
print(f"[app] ASR transcription failed: {type(exc).__name__}: {exc}")
|
| 267 |
+
asr_metadata = {
|
| 268 |
+
"model": None,
|
| 269 |
+
"language": language,
|
| 270 |
+
"status": "error",
|
| 271 |
+
"error": f"{type(exc).__name__}: {exc}",
|
| 272 |
+
}
|
| 273 |
+
|
| 274 |
+
if not transcript or not transcript.strip():
|
| 275 |
+
return (
|
| 276 |
+
"β οΈ No transcript available β record audio or type a note first."
|
| 277 |
+
)
|
| 278 |
+
|
| 279 |
+
# ββ 2. Build journal entry ββββββββββββββββββββββββββββββββββββββββ
|
| 280 |
entry = create_journal_entry(
|
| 281 |
transcript=transcript,
|
| 282 |
session_id=session_id,
|
| 283 |
team_id=team_id,
|
| 284 |
task_id=task_id or None,
|
| 285 |
location_note=location_note,
|
| 286 |
+
audio_ref=audio_ref,
|
| 287 |
+
asr_metadata=asr_metadata,
|
| 288 |
+
transcript_source=transcript_source,
|
| 289 |
)
|
| 290 |
|
| 291 |
# Summarize
|
|
|
|
| 303 |
"mood": entry["mood"],
|
| 304 |
"story_value": summary["story_value"],
|
| 305 |
"summary": summary["moment_summary"],
|
| 306 |
+
"transcript_source": transcript_source,
|
| 307 |
+
"asr_status": asr_metadata.get("status") if asr_metadata else None,
|
| 308 |
+
"asr_model": asr_metadata.get("model") if asr_metadata else None,
|
| 309 |
}, team_id=team_id)
|
| 310 |
SESSION_STORE[session_id]["events"].append(ev)
|
| 311 |
SESSION_STORE[session_id]["journals"].append(entry)
|
| 312 |
|
| 313 |
# Build display text
|
| 314 |
+
source_label = {
|
| 315 |
+
"asr": "ποΈ (transcribed)",
|
| 316 |
+
"hybrid": "ποΈβοΈ (transcribed, edited)",
|
| 317 |
+
"typed": "β¨οΈ (typed)",
|
| 318 |
+
}.get(transcript_source, transcript_source)
|
| 319 |
+
asr_line = ""
|
| 320 |
+
if asr_metadata and asr_metadata.get("status") != "ok":
|
| 321 |
+
asr_line = f"- ASR status: **{asr_metadata.get('status')}** β {asr_metadata.get('error') or ''}\n"
|
| 322 |
display = (
|
| 323 |
+
f"ποΈ **Journal recorded!** {source_label}\n"
|
| 324 |
f"- Mood: *{entry['mood']}*\n"
|
| 325 |
f"- Story value: **{summary['story_value']}**\n"
|
| 326 |
f"- Tags: {', '.join(summary['tags'])}\n"
|
| 327 |
+
f"{asr_line}"
|
| 328 |
f"- Summary: {summary['moment_summary']}"
|
| 329 |
)
|
| 330 |
return display
|
|
|
|
| 1075 |
task_feedback = gr.Markdown(value="")
|
| 1076 |
|
| 1077 |
gr.HTML('<div class="section-label" style="margin-top:20px;">Voice Journal</div>')
|
| 1078 |
+
journal_audio = gr.Audio(
|
| 1079 |
+
label="ποΈ Record voice (auto-transcribed via Cohere ASR)",
|
| 1080 |
+
sources=["microphone", "upload"],
|
| 1081 |
+
type="filepath",
|
| 1082 |
+
)
|
| 1083 |
journal_transcript = gr.Textbox(
|
| 1084 |
+
label="Transcript (edit the ASR result or type your own)",
|
| 1085 |
+
lines=3,
|
| 1086 |
placeholder="We just found the mural near the canal β it was incredible!",
|
| 1087 |
)
|
| 1088 |
with gr.Row():
|
| 1089 |
journal_task_id = gr.Textbox(label="Task ID (opt.)", placeholder="t1")
|
| 1090 |
journal_location = gr.Textbox(label="Location note", placeholder="Rue de Rivoli")
|
| 1091 |
+
journal_language = gr.Dropdown(
|
| 1092 |
+
label="Language", value="en",
|
| 1093 |
+
choices=["en", "fr", "de", "it", "es", "pt", "el", "nl", "pl",
|
| 1094 |
+
"zh", "ja", "ko", "vi", "ar"],
|
| 1095 |
+
)
|
| 1096 |
journal_btn = gr.Button("ποΈ Record Journal", variant="secondary")
|
| 1097 |
journal_output = gr.Markdown(value="")
|
| 1098 |
|
|
|
|
| 1304 |
|
| 1305 |
journal_btn.click(
|
| 1306 |
fn=record_journal,
|
| 1307 |
+
inputs=[current_session, journal_transcript, journal_task_id, journal_location,
|
| 1308 |
+
current_team, journal_audio, journal_language],
|
| 1309 |
outputs=[journal_output],
|
| 1310 |
)
|
| 1311 |
photo_btn.click(
|
app/data/sessions_store.json
CHANGED
|
@@ -1 +1 @@
|
|
| 1 |
-
{"codes": {"75ZKFE": "9fcb0953-4e36-4485-9bbe-fad577de31fa", "JWR5RV": "6b1af473-b00b-44ae-bffc-7b8e93616767"}, "sessions": {"9fcb0953-4e36-4485-9bbe-fad577de31fa": {"config": {"game_type": "scavenger_hunt", "city": "Paris", "area": "Le Marais", "location_type": "mixed", "duration_minutes": 60, "num_players": 2, "difficulty": "easy", "age_group": "adults", "energy_level": "low", "photo_enabled": true}, "game": {"game_id": "mock-8506e14a", "title": "Scavenger_Hunt in Le Marais", "theme": "easy adventure", "setup": {"city": "Paris", "area": "Le Marais", "meeting_point": "Main entrance of Le Marais", "duration_minutes": 60, "num_players": 2}, "rules": ["Complete as many tasks as possible within 60 minutes", "Take photos or notes as proof of completion", "Stay within the designated area at all times", "No entering private buildings or restricted areas", "This game is suitable for adults"], "tasks": [{"task_id": "t1", "title": "Task 1: Explore the main square", "description": "Find and document something interesting in the main square", "location_hint": "Navigate to the main square and look for distinctive features", "points": 15, "time_limit_minutes": 8, "proof_type": "photo", "hint": "Look for signs or landmarks in the main square", "safety_note": "Stay on public paths and avoid restricted areas"}, {"task_id": "t2", "title": "Task 2: Explore the city center", "description": "Find and document something interesting in the city center", "location_hint": "Navigate to the city center and look for distinctive features", "points": 20, "time_limit_minutes": 10, "proof_type": "observation", "hint": "Look for signs or landmarks in the city center", "safety_note": "Stay on public paths and avoid restricted areas"}, {"task_id": "t3", "title": "Task 3: Explore the park area", "description": "Find and document something interesting in the park area", "location_hint": "Navigate to the park area and look for distinctive features", "points": 25, "time_limit_minutes": 12, "proof_type": "text", "hint": "Look for signs or landmarks in the park area", "safety_note": "Stay on public paths and avoid restricted areas"}, {"task_id": "t4", "title": "Task 4: Explore the landmark district", "description": "Find and document something interesting in the landmark district", "location_hint": "Navigate to the landmark district and look for distinctive features", "points": 30, "time_limit_minutes": 14, "proof_type": "photo", "hint": "Look for signs or landmarks in the landmark district", "safety_note": "Stay on public paths and avoid restricted areas"}], "global_hints": ["Explore systematically from the meeting point outward", "Ask locals for directions if needed", "Time management is key - don't spend too long on any single task"], "score_rules": ["Each task completed: full points", "Early completion: +1 bonus point per minute under limit", "Hints used: -5 points per hint", "Late arrival at meeting point: -10 points per minute"], "tie_breaker": "Winner is the player with the most points when time expires. Ties broken by earliest completion time.", "safety": {"allowed_zone": "Le Marais", "forbidden_behaviors": ["Entering buildings without permission", "Crossing busy streets recklessly", "Approaching strangers", "Leaving the designated area"], "adult_supervision": false, "stop_conditions": ["If a player feels unsafe, the game stops immediately", "If weather becomes severe, relocate to shelter", "If anyone is injured, call emergency services"]}, "story_seed": {"tone": "playful", "motifs": ["exploration", "discovery", "teamwork"], "recap_style": "episode_recap"}}, "adventure_code": "75ZKFE", "teams": ["team-a", "team-b"], "players": [{"name": "Player 1", "team_id": "team-a"}, {"name": "play 2", "team_id": "team-b"}, {"name": "hola", "team_id": "team-b"}], "photos": [{"photo_id": "photo-71c44ee8", "photo_name": "", "caption": "", "task_id": "t1"}]}, "6b1af473-b00b-44ae-bffc-7b8e93616767": {"config": {"game_type": "scavenger_hunt", "city": "Paris", "area": "Le Marais", "location_type": "mixed", "duration_minutes": 60, "num_players": 2, "difficulty": "easy", "age_group": "adults", "energy_level": "low", "photo_enabled": true}, "game": {"game_id": "mock-6d6ab9cf", "title": "Scavenger_Hunt in Le Marais", "theme": "easy adventure", "setup": {"city": "Paris", "area": "Le Marais", "meeting_point": "Main entrance of Le Marais", "duration_minutes": 60, "num_players": 2}, "rules": ["Complete as many tasks as possible within 60 minutes", "Take photos or notes as proof of completion", "Stay within the designated area at all times", "No entering private buildings or restricted areas", "This game is suitable for adults"], "tasks": [{"task_id": "t1", "title": "Task 1: Explore the main square", "description": "Find and document something interesting in the main square", "location_hint": "Navigate to the main square and look for distinctive features", "points": 15, "time_limit_minutes": 8, "proof_type": "photo", "hint": "Look for signs or landmarks in the main square", "safety_note": "Stay on public paths and avoid restricted areas"}, {"task_id": "t2", "title": "Task 2: Explore the city center", "description": "Find and document something interesting in the city center", "location_hint": "Navigate to the city center and look for distinctive features", "points": 20, "time_limit_minutes": 10, "proof_type": "observation", "hint": "Look for signs or landmarks in the city center", "safety_note": "Stay on public paths and avoid restricted areas"}, {"task_id": "t3", "title": "Task 3: Explore the park area", "description": "Find and document something interesting in the park area", "location_hint": "Navigate to the park area and look for distinctive features", "points": 25, "time_limit_minutes": 12, "proof_type": "text", "hint": "Look for signs or landmarks in the park area", "safety_note": "Stay on public paths and avoid restricted areas"}, {"task_id": "t4", "title": "Task 4: Explore the landmark district", "description": "Find and document something interesting in the landmark district", "location_hint": "Navigate to the landmark district and look for distinctive features", "points": 30, "time_limit_minutes": 14, "proof_type": "photo", "hint": "Look for signs or landmarks in the landmark district", "safety_note": "Stay on public paths and avoid restricted areas"}], "global_hints": ["Explore systematically from the meeting point outward", "Ask locals for directions if needed", "Time management is key - don't spend too long on any single task"], "score_rules": ["Each task completed: full points", "Early completion: +1 bonus point per minute under limit", "Hints used: -5 points per hint", "Late arrival at meeting point: -10 points per minute"], "tie_breaker": "Winner is the player with the most points when time expires. Ties broken by earliest completion time.", "safety": {"allowed_zone": "Le Marais", "forbidden_behaviors": ["Entering buildings without permission", "Crossing busy streets recklessly", "Approaching strangers", "Leaving the designated area"], "adult_supervision": false, "stop_conditions": ["If a player feels unsafe, the game stops immediately", "If weather becomes severe, relocate to shelter", "If anyone is injured, call emergency services"]}, "story_seed": {"tone": "playful", "motifs": ["exploration", "discovery", "teamwork"], "recap_style": "episode_recap"}}, "adventure_code": "JWR5RV", "teams": ["team-a", "team-b"], "players": [{"name": "Player 1", "team_id": "team-a"}], "photos": []}}}
|
|
|
|
| 1 |
+
{"codes": {"75ZKFE": "9fcb0953-4e36-4485-9bbe-fad577de31fa", "JWR5RV": "6b1af473-b00b-44ae-bffc-7b8e93616767", "NKPZKT": "32e1bbe3-50a1-4e43-a7c2-8ecf3f781550", "NXWHQK": "4b03fbfc-4326-4325-8975-e4216c44e03e"}, "sessions": {"9fcb0953-4e36-4485-9bbe-fad577de31fa": {"config": {"game_type": "scavenger_hunt", "city": "Paris", "area": "Le Marais", "location_type": "mixed", "duration_minutes": 60, "num_players": 2, "difficulty": "easy", "age_group": "adults", "energy_level": "low", "photo_enabled": true}, "game": {"game_id": "mock-8506e14a", "title": "Scavenger_Hunt in Le Marais", "theme": "easy adventure", "setup": {"city": "Paris", "area": "Le Marais", "meeting_point": "Main entrance of Le Marais", "duration_minutes": 60, "num_players": 2}, "rules": ["Complete as many tasks as possible within 60 minutes", "Take photos or notes as proof of completion", "Stay within the designated area at all times", "No entering private buildings or restricted areas", "This game is suitable for adults"], "tasks": [{"task_id": "t1", "title": "Task 1: Explore the main square", "description": "Find and document something interesting in the main square", "location_hint": "Navigate to the main square and look for distinctive features", "points": 15, "time_limit_minutes": 8, "proof_type": "photo", "hint": "Look for signs or landmarks in the main square", "safety_note": "Stay on public paths and avoid restricted areas"}, {"task_id": "t2", "title": "Task 2: Explore the city center", "description": "Find and document something interesting in the city center", "location_hint": "Navigate to the city center and look for distinctive features", "points": 20, "time_limit_minutes": 10, "proof_type": "observation", "hint": "Look for signs or landmarks in the city center", "safety_note": "Stay on public paths and avoid restricted areas"}, {"task_id": "t3", "title": "Task 3: Explore the park area", "description": "Find and document something interesting in the park area", "location_hint": "Navigate to the park area and look for distinctive features", "points": 25, "time_limit_minutes": 12, "proof_type": "text", "hint": "Look for signs or landmarks in the park area", "safety_note": "Stay on public paths and avoid restricted areas"}, {"task_id": "t4", "title": "Task 4: Explore the landmark district", "description": "Find and document something interesting in the landmark district", "location_hint": "Navigate to the landmark district and look for distinctive features", "points": 30, "time_limit_minutes": 14, "proof_type": "photo", "hint": "Look for signs or landmarks in the landmark district", "safety_note": "Stay on public paths and avoid restricted areas"}], "global_hints": ["Explore systematically from the meeting point outward", "Ask locals for directions if needed", "Time management is key - don't spend too long on any single task"], "score_rules": ["Each task completed: full points", "Early completion: +1 bonus point per minute under limit", "Hints used: -5 points per hint", "Late arrival at meeting point: -10 points per minute"], "tie_breaker": "Winner is the player with the most points when time expires. Ties broken by earliest completion time.", "safety": {"allowed_zone": "Le Marais", "forbidden_behaviors": ["Entering buildings without permission", "Crossing busy streets recklessly", "Approaching strangers", "Leaving the designated area"], "adult_supervision": false, "stop_conditions": ["If a player feels unsafe, the game stops immediately", "If weather becomes severe, relocate to shelter", "If anyone is injured, call emergency services"]}, "story_seed": {"tone": "playful", "motifs": ["exploration", "discovery", "teamwork"], "recap_style": "episode_recap"}}, "adventure_code": "75ZKFE", "teams": ["team-a", "team-b"], "players": [{"name": "Player 1", "team_id": "team-a"}, {"name": "play 2", "team_id": "team-b"}, {"name": "hola", "team_id": "team-b"}], "photos": [{"photo_id": "photo-71c44ee8", "photo_name": "", "caption": "", "task_id": "t1"}]}, "6b1af473-b00b-44ae-bffc-7b8e93616767": {"config": {"game_type": "scavenger_hunt", "city": "Paris", "area": "Le Marais", "location_type": "mixed", "duration_minutes": 60, "num_players": 2, "difficulty": "easy", "age_group": "adults", "energy_level": "low", "photo_enabled": true}, "game": {"game_id": "mock-6d6ab9cf", "title": "Scavenger_Hunt in Le Marais", "theme": "easy adventure", "setup": {"city": "Paris", "area": "Le Marais", "meeting_point": "Main entrance of Le Marais", "duration_minutes": 60, "num_players": 2}, "rules": ["Complete as many tasks as possible within 60 minutes", "Take photos or notes as proof of completion", "Stay within the designated area at all times", "No entering private buildings or restricted areas", "This game is suitable for adults"], "tasks": [{"task_id": "t1", "title": "Task 1: Explore the main square", "description": "Find and document something interesting in the main square", "location_hint": "Navigate to the main square and look for distinctive features", "points": 15, "time_limit_minutes": 8, "proof_type": "photo", "hint": "Look for signs or landmarks in the main square", "safety_note": "Stay on public paths and avoid restricted areas"}, {"task_id": "t2", "title": "Task 2: Explore the city center", "description": "Find and document something interesting in the city center", "location_hint": "Navigate to the city center and look for distinctive features", "points": 20, "time_limit_minutes": 10, "proof_type": "observation", "hint": "Look for signs or landmarks in the city center", "safety_note": "Stay on public paths and avoid restricted areas"}, {"task_id": "t3", "title": "Task 3: Explore the park area", "description": "Find and document something interesting in the park area", "location_hint": "Navigate to the park area and look for distinctive features", "points": 25, "time_limit_minutes": 12, "proof_type": "text", "hint": "Look for signs or landmarks in the park area", "safety_note": "Stay on public paths and avoid restricted areas"}, {"task_id": "t4", "title": "Task 4: Explore the landmark district", "description": "Find and document something interesting in the landmark district", "location_hint": "Navigate to the landmark district and look for distinctive features", "points": 30, "time_limit_minutes": 14, "proof_type": "photo", "hint": "Look for signs or landmarks in the landmark district", "safety_note": "Stay on public paths and avoid restricted areas"}], "global_hints": ["Explore systematically from the meeting point outward", "Ask locals for directions if needed", "Time management is key - don't spend too long on any single task"], "score_rules": ["Each task completed: full points", "Early completion: +1 bonus point per minute under limit", "Hints used: -5 points per hint", "Late arrival at meeting point: -10 points per minute"], "tie_breaker": "Winner is the player with the most points when time expires. Ties broken by earliest completion time.", "safety": {"allowed_zone": "Le Marais", "forbidden_behaviors": ["Entering buildings without permission", "Crossing busy streets recklessly", "Approaching strangers", "Leaving the designated area"], "adult_supervision": false, "stop_conditions": ["If a player feels unsafe, the game stops immediately", "If weather becomes severe, relocate to shelter", "If anyone is injured, call emergency services"]}, "story_seed": {"tone": "playful", "motifs": ["exploration", "discovery", "teamwork"], "recap_style": "episode_recap"}}, "adventure_code": "JWR5RV", "teams": ["team-a", "team-b"], "players": [{"name": "Player 1", "team_id": "team-a"}], "photos": []}, "32e1bbe3-50a1-4e43-a7c2-8ecf3f781550": {"config": {"game_type": "scavenger_hunt", "city": "Paris", "area": "Le Marais", "location_type": "mixed", "duration_minutes": 60, "num_players": 4, "difficulty": "medium", "age_group": "adults", "energy_level": "medium", "photo_enabled": true}, "game": {"game_id": "game_042", "title": "Le Marais Mystery Hunt", "theme": "Scavenger Hunt", "setup": {"city": "Paris", "area": "Le Marais", "meeting_point": "Place des Vosges", "duration_minutes": 60, "num_players": 4}, "rules": ["Stay within Le Marais boundaries", "No entering buildings or private property", "No proximity to water, traffic, or rail lines", "No interacting with strangers", "All tasks must be observed, no purchases"], "tasks": [{"task_id": "t1", "title": "Historic Fountain", "description": "Observe the historic fountain at Place des Vosges.", "location_hint": "Place des Vosges", "points": 20, "time_limit_minutes": null, "proof_type": "observation", "hint": "Look for the historic fountain", "safety_note": "Only observe, no touching or entering"}, {"task_id": "t2", "title": "Colorful Door", "description": "Observe the colorful door on Rue des Rosiers.", "location_hint": "Rue des Rosiers", "points": 15, "time_limit_minutes": null, "proof_type": "observation", "hint": "Look for the colorful door", "safety_note": "Only observe, no touching or entering"}, {"task_id": "t3", "title": "Task 3", "description": "Explore and discover something interesting in the area", "location_hint": "Look around the central area", "points": 25, "time_limit_minutes": null, "proof_type": "text", "hint": "Check visible landmarks and signs for clues", "safety_note": "Stay on public paths and respect local rules"}], "global_hints": ["Le Marais is a historic district", "All tasks are observation-based"], "score_rules": ["Higher points win", "Tie broken by total points"], "tie_breaker": "Highest total points wins", "safety": {"allowed_zone": "Le Marais public streets and parks", "forbidden_behaviors": ["Entering buildings or private property", "Proximity to water, traffic, or rail lines", "Interacting with strangers", "Purchasing items"], "adult_supervision": false, "stop_conditions": ["All tasks completed", "Time limit 60 minutes reached", "Any rule violation"]}, "story_seed": {"tone": "playful", "motifs": ["historic architecture", "urban exploration"], "recap_style": "Summarize each player's observations and points"}}, "adventure_code": "NKPZKT", "teams": ["team-a", "team-b"], "players": [{"name": "Nithin", "team_id": "team-a"}], "photos": []}, "4b03fbfc-4326-4325-8975-e4216c44e03e": {"config": {"game_type": "scavenger_hunt", "city": "Paris", "area": "Le Marais", "location_type": "mixed", "duration_minutes": 60, "num_players": 4, "difficulty": "medium", "age_group": "adults", "energy_level": "medium", "photo_enabled": true}, "game": {"game_id": "game_042", "title": "Le Marais Mystery Hunt", "theme": "Observational scavenger hunt", "setup": {"city": "Paris", "area": "Le Marais", "meeting_point": "Place des Vosges", "duration_minutes": 60, "num_players": 4}, "rules": ["Stay within Le Marais boundaries", "No entering buildings or private property", "No crossing traffic or water", "No interacting with strangers", "No purchases", "All tasks are observation only"], "tasks": [{"task_id": "t1", "title": "Vosges View", "description": "Find the historic square Place des Vosges and note its symmetrical layout.", "location_hint": "Place des Vosges", "points": 20, "time_limit_minutes": null, "proof_type": "observation", "hint": "Look for the central fountain", "safety_note": "Stay on pavement, no building entry"}, {"task_id": "t2", "title": "Rue des Rosiers", "description": "Observe the vibrant street market stalls and colorful awnings.", "location_hint": "Rue des Rosiers", "points": 15, "time_limit_minutes": null, "proof_type": "observation", "hint": "Notice the flower displays", "safety_note": "No entering stalls"}, {"task_id": "t3", "title": "H\u00f4tel de Ville View", "description": "Observe the fa\u00e7ade of the H\u00f4tel de Ville from the street.", "location_hint": "Place des Vosges near H\u00f4tel de Ville", "points": 20, "time_limit_minutes": null, "proof_type": "observation", "hint": "Look for the clock tower", "safety_note": "No entering building"}], "global_hints": ["Stay within Le Marais", "Use observation only", "No purchases"], "score_rules": ["Highest points wins", "All tasks completed for bonus"], "tie_breaker": "First player to reach 50 points wins", "safety": {"allowed_zone": "Le Marais streets and public squares", "forbidden_behaviors": ["Entering buildings", "Crossing traffic", "Approaching water", "Interacting with strangers", "Purchasing items"], "adult_supervision": true, "stop_conditions": ["Time limit 60 minutes", "Safety breach", "All tasks completed"]}, "story_seed": {"tone": "playful", "motifs": ["hidden clues", "Parisian charm", "observation"], "recap_style": "Gather at the meeting point for a recap of findings."}}, "adventure_code": "NXWHQK", "teams": ["team-a", "team-b"], "players": [{"name": "Nithin", "team_id": "team-a"}], "photos": []}}}
|
app/schemas/journal_schema.json
CHANGED
|
@@ -52,6 +52,38 @@
|
|
| 52 |
"type": "string"
|
| 53 |
},
|
| 54 |
"description": "References to associated photos"
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 55 |
}
|
| 56 |
}
|
| 57 |
}
|
|
|
|
| 52 |
"type": "string"
|
| 53 |
},
|
| 54 |
"description": "References to associated photos"
|
| 55 |
+
},
|
| 56 |
+
"transcript_source": {
|
| 57 |
+
"type": "string",
|
| 58 |
+
"enum": ["typed", "asr", "hybrid"],
|
| 59 |
+
"description": "Where the transcript text came from: typed by the user, produced by the ASR service, or hybrid (ASR + manual edits)."
|
| 60 |
+
},
|
| 61 |
+
"audio_ref": {
|
| 62 |
+
"type": "string",
|
| 63 |
+
"description": "Path or URL of the recorded audio clip associated with this entry."
|
| 64 |
+
},
|
| 65 |
+
"asr": {
|
| 66 |
+
"type": "object",
|
| 67 |
+
"description": "Metadata describing the ASR pass that produced the transcript (when transcript_source is 'asr' or 'hybrid').",
|
| 68 |
+
"properties": {
|
| 69 |
+
"model": {
|
| 70 |
+
"type": "string",
|
| 71 |
+
"description": "ASR model identifier (e.g. CohereLabs/cohere-transcribe-03-2026)."
|
| 72 |
+
},
|
| 73 |
+
"language": {
|
| 74 |
+
"type": "string",
|
| 75 |
+
"description": "Language code used for transcription."
|
| 76 |
+
},
|
| 77 |
+
"status": {
|
| 78 |
+
"type": "string",
|
| 79 |
+
"enum": ["ok", "skipped", "error"],
|
| 80 |
+
"description": "Outcome of the ASR pass."
|
| 81 |
+
},
|
| 82 |
+
"error": {
|
| 83 |
+
"type": ["string", "null"],
|
| 84 |
+
"description": "Error message if status is 'error' or 'skipped'."
|
| 85 |
+
}
|
| 86 |
+
}
|
| 87 |
}
|
| 88 |
}
|
| 89 |
}
|
app/services/asr.py
ADDED
|
@@ -0,0 +1,269 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Automatic Speech Recognition (ASR) module.
|
| 2 |
+
|
| 3 |
+
Default model: ``CohereLabs/cohere-transcribe-03-2026`` (Apache 2.0, 2B
|
| 4 |
+
parameters, 14 languages).
|
| 5 |
+
|
| 6 |
+
Uses the official π€ Transformers integration:
|
| 7 |
+
|
| 8 |
+
- transformers >= 5.4.0
|
| 9 |
+
- torch
|
| 10 |
+
- huggingface_hub
|
| 11 |
+
- soundfile, librosa
|
| 12 |
+
- sentencepiece, protobuf
|
| 13 |
+
|
| 14 |
+
Audio is auto-resampled to 16 kHz and downmixed to mono by the
|
| 15 |
+
processor. The model supports 14 languages (English, French, German,
|
| 16 |
+
Italian, Spanish, Portuguese, Greek, Dutch, Polish, Mandarin, Japanese,
|
| 17 |
+
Korean, Vietnamese, Arabic) but does not auto-detect language, so
|
| 18 |
+
callers should pass an explicit language code.
|
| 19 |
+
|
| 20 |
+
Like the other model services, the model is **lazy-loaded** so it does
|
| 21 |
+
not block app startup. Setting the environment variable
|
| 22 |
+
``CITYQUEST_SKIP_MODEL=1`` (or ``CITYQUEST_FAST_TEST=1``) will make all
|
| 23 |
+
ASR calls return an empty string with a warning so the rest of the
|
| 24 |
+
pipeline (typed journal fallback, scoring, recap) keeps working.
|
| 25 |
+
"""
|
| 26 |
+
|
| 27 |
+
from __future__ import annotations
|
| 28 |
+
|
| 29 |
+
import os
|
| 30 |
+
from pathlib import Path
|
| 31 |
+
from typing import Optional
|
| 32 |
+
|
| 33 |
+
|
| 34 |
+
# ββ Configuration βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
|
| 35 |
+
MODEL_ID = "CohereLabs/cohere-transcribe-03-2026"
|
| 36 |
+
SUPPORTED_LANGUAGES = [
|
| 37 |
+
"en", "fr", "de", "it", "es", "pt", "el", "nl", "pl",
|
| 38 |
+
"zh", "ja", "ko", "vi", "ar",
|
| 39 |
+
]
|
| 40 |
+
DEFAULT_LANGUAGE = "en"
|
| 41 |
+
MAX_NEW_TOKENS = 256
|
| 42 |
+
SAMPLING_RATE = 16_000
|
| 43 |
+
|
| 44 |
+
# Skipped under CITYQUEST_SKIP_MODEL (model download) and CITYQUEST_FAST_TEST
|
| 45 |
+
# (used by `test_end_to_end.py`).
|
| 46 |
+
_skip_env_var = "CITYQUEST_SKIP_MODEL"
|
| 47 |
+
_fast_test_env_var = "CITYQUEST_FAST_TEST"
|
| 48 |
+
|
| 49 |
+
|
| 50 |
+
# ββ Lazy model state ββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
|
| 51 |
+
_processor = None
|
| 52 |
+
_model = None
|
| 53 |
+
_model_loaded = False
|
| 54 |
+
_load_attempted = False
|
| 55 |
+
|
| 56 |
+
|
| 57 |
+
def _is_skipped() -> bool:
|
| 58 |
+
"""True if ASR model loading/inference should be short-circuited."""
|
| 59 |
+
return bool(os.environ.get(_skip_env_var) or os.environ.get(_fast_test_env_var))
|
| 60 |
+
|
| 61 |
+
|
| 62 |
+
def _load_model() -> bool:
|
| 63 |
+
"""Lazy-load the Cohere Transcribe model + processor on first call.
|
| 64 |
+
|
| 65 |
+
Returns:
|
| 66 |
+
``True`` if the model is ready, ``False`` if it was skipped or
|
| 67 |
+
failed to load.
|
| 68 |
+
"""
|
| 69 |
+
global _processor, _model, _model_loaded, _load_attempted
|
| 70 |
+
if _model_loaded:
|
| 71 |
+
return _processor is not None and _model is not None
|
| 72 |
+
_load_attempted = True
|
| 73 |
+
|
| 74 |
+
if _is_skipped():
|
| 75 |
+
print("[asr] skip env var set β ASR model not loaded")
|
| 76 |
+
return False
|
| 77 |
+
|
| 78 |
+
try:
|
| 79 |
+
import torch
|
| 80 |
+
from transformers import AutoProcessor, CohereAsrForConditionalGeneration
|
| 81 |
+
from huggingface_hub import login
|
| 82 |
+
|
| 83 |
+
# Try logging in if a token is provided in env (gated model).
|
| 84 |
+
token = os.environ.get("HF_TOKEN") or os.environ.get("HUGGINGFACE_HUB_TOKEN")
|
| 85 |
+
if token:
|
| 86 |
+
try:
|
| 87 |
+
login(token=token, add_to_git_credential=False)
|
| 88 |
+
except Exception as exc:
|
| 89 |
+
print(f"[asr] HF login skipped: {exc}")
|
| 90 |
+
|
| 91 |
+
device = "cuda" if torch.cuda.is_available() else "cpu"
|
| 92 |
+
dtype = torch.bfloat16 if device == "cuda" else torch.float32
|
| 93 |
+
|
| 94 |
+
print(f"[asr] Loading {MODEL_ID} on {device} ({dtype}) β¦")
|
| 95 |
+
_processor = AutoProcessor.from_pretrained(MODEL_ID)
|
| 96 |
+
_model = CohereAsrForConditionalGeneration.from_pretrained(
|
| 97 |
+
MODEL_ID, torch_dtype=dtype, device_map="auto",
|
| 98 |
+
)
|
| 99 |
+
_model.eval()
|
| 100 |
+
_model_loaded = True
|
| 101 |
+
print("[asr] Model ready")
|
| 102 |
+
return True
|
| 103 |
+
|
| 104 |
+
except ImportError as exc:
|
| 105 |
+
print(
|
| 106 |
+
f"[asr] Required packages not installed: {exc}. "
|
| 107 |
+
"Install with: pip install 'transformers>=5.4.0' torch "
|
| 108 |
+
"huggingface_hub soundfile librosa sentencepiece protobuf"
|
| 109 |
+
)
|
| 110 |
+
except Exception as exc:
|
| 111 |
+
print(f"[asr] Model load failed: {type(exc).__name__}: {exc}")
|
| 112 |
+
return False
|
| 113 |
+
|
| 114 |
+
|
| 115 |
+
def _load_audio(audio_path: str) -> Optional[list[float]]:
|
| 116 |
+
"""Load an audio file as a mono float32 waveform at 16 kHz.
|
| 117 |
+
|
| 118 |
+
Tries multiple loading strategies in order:
|
| 119 |
+
1. ``transformers.audio_utils.load_audio`` (preferred, handles
|
| 120 |
+
resampling + mono conversion).
|
| 121 |
+
2. ``soundfile`` fallback.
|
| 122 |
+
"""
|
| 123 |
+
if not Path(audio_path).exists():
|
| 124 |
+
print(f"[asr] Audio file not found: {audio_path}")
|
| 125 |
+
return None
|
| 126 |
+
|
| 127 |
+
try:
|
| 128 |
+
from transformers.audio_utils import load_audio
|
| 129 |
+
|
| 130 |
+
waveform = load_audio(audio_path, sampling_rate=SAMPLING_RATE)
|
| 131 |
+
return list(waveform)
|
| 132 |
+
except ImportError:
|
| 133 |
+
pass
|
| 134 |
+
except Exception as exc:
|
| 135 |
+
print(f"[asr] transformers audio load failed: {exc}")
|
| 136 |
+
|
| 137 |
+
# Fallback: soundfile + numpy
|
| 138 |
+
try:
|
| 139 |
+
import numpy as np
|
| 140 |
+
import soundfile as sf
|
| 141 |
+
|
| 142 |
+
data, sr = sf.read(audio_path, dtype="float32", always_2d=False)
|
| 143 |
+
if data.ndim > 1:
|
| 144 |
+
data = data.mean(axis=1)
|
| 145 |
+
if sr != SAMPLING_RATE:
|
| 146 |
+
# simple linear resample β good enough for short voice notes
|
| 147 |
+
ratio = SAMPLING_RATE / sr
|
| 148 |
+
new_len = int(len(data) * ratio)
|
| 149 |
+
data = np.interp(
|
| 150 |
+
np.linspace(0, len(data), new_len, endpoint=False),
|
| 151 |
+
np.arange(len(data)),
|
| 152 |
+
data,
|
| 153 |
+
).astype("float32")
|
| 154 |
+
return list(data)
|
| 155 |
+
except ImportError:
|
| 156 |
+
print(
|
| 157 |
+
"[asr] soundfile not installed β install with: "
|
| 158 |
+
"pip install soundfile librosa"
|
| 159 |
+
)
|
| 160 |
+
except Exception as exc:
|
| 161 |
+
print(f"[asr] soundfile load failed: {exc}")
|
| 162 |
+
return None
|
| 163 |
+
|
| 164 |
+
|
| 165 |
+
# ββ Public API ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
|
| 166 |
+
|
| 167 |
+
def transcribe(
|
| 168 |
+
audio_path: str,
|
| 169 |
+
language: str = DEFAULT_LANGUAGE,
|
| 170 |
+
max_new_tokens: int = MAX_NEW_TOKENS,
|
| 171 |
+
) -> dict:
|
| 172 |
+
"""Transcribe a voice journal audio clip to text.
|
| 173 |
+
|
| 174 |
+
Args:
|
| 175 |
+
audio_path: Path to an audio file (wav/mp3/m4a/ogg/webm).
|
| 176 |
+
language: One of ``SUPPORTED_LANGUAGES``. Defaults to ``"en"``.
|
| 177 |
+
max_new_tokens: Maximum number of tokens to generate.
|
| 178 |
+
|
| 179 |
+
Returns:
|
| 180 |
+
Dict with keys ``transcript`` (str), ``language`` (str),
|
| 181 |
+
``model`` (str), ``status`` (str β ``"ok"`` | ``"skipped"`` |
|
| 182 |
+
``"error"``), and ``error`` (str | None). On non-ok statuses
|
| 183 |
+
``transcript`` is ``""``.
|
| 184 |
+
"""
|
| 185 |
+
if language not in SUPPORTED_LANGUAGES:
|
| 186 |
+
print(
|
| 187 |
+
f"[asr] Language '{language}' not in supported set β "
|
| 188 |
+
f"falling back to '{DEFAULT_LANGUAGE}'."
|
| 189 |
+
)
|
| 190 |
+
language = DEFAULT_LANGUAGE
|
| 191 |
+
|
| 192 |
+
if _is_skipped():
|
| 193 |
+
return {
|
| 194 |
+
"transcript": "",
|
| 195 |
+
"language": language,
|
| 196 |
+
"model": MODEL_ID,
|
| 197 |
+
"status": "skipped",
|
| 198 |
+
"error": "ASR model skipped (CITYQUEST_SKIP_MODEL/CITYQUEST_FAST_TEST).",
|
| 199 |
+
}
|
| 200 |
+
|
| 201 |
+
if not _load_model():
|
| 202 |
+
return {
|
| 203 |
+
"transcript": "",
|
| 204 |
+
"language": language,
|
| 205 |
+
"model": MODEL_ID,
|
| 206 |
+
"status": "error",
|
| 207 |
+
"error": "ASR model unavailable.",
|
| 208 |
+
}
|
| 209 |
+
|
| 210 |
+
waveform = _load_audio(audio_path)
|
| 211 |
+
if waveform is None:
|
| 212 |
+
return {
|
| 213 |
+
"transcript": "",
|
| 214 |
+
"language": language,
|
| 215 |
+
"model": MODEL_ID,
|
| 216 |
+
"status": "error",
|
| 217 |
+
"error": f"Could not load audio at {audio_path}.",
|
| 218 |
+
}
|
| 219 |
+
|
| 220 |
+
try:
|
| 221 |
+
import torch
|
| 222 |
+
|
| 223 |
+
inputs = _processor(
|
| 224 |
+
waveform,
|
| 225 |
+
sampling_rate=SAMPLING_RATE,
|
| 226 |
+
return_tensors="pt",
|
| 227 |
+
language=language,
|
| 228 |
+
)
|
| 229 |
+
inputs = {
|
| 230 |
+
k: v.to(_model.device, dtype=_model.dtype) if hasattr(v, "to") else v
|
| 231 |
+
for k, v in inputs.items()
|
| 232 |
+
}
|
| 233 |
+
|
| 234 |
+
with torch.inference_mode():
|
| 235 |
+
outputs = _model.generate(
|
| 236 |
+
**inputs,
|
| 237 |
+
max_new_tokens=max_new_tokens,
|
| 238 |
+
)
|
| 239 |
+
text = _processor.decode(outputs[0], skip_special_tokens=True).strip()
|
| 240 |
+
|
| 241 |
+
return {
|
| 242 |
+
"transcript": text,
|
| 243 |
+
"language": language,
|
| 244 |
+
"model": MODEL_ID,
|
| 245 |
+
"status": "ok",
|
| 246 |
+
"error": None,
|
| 247 |
+
}
|
| 248 |
+
|
| 249 |
+
except Exception as exc:
|
| 250 |
+
print(f"[asr] Inference failed: {type(exc).__name__}: {exc}")
|
| 251 |
+
return {
|
| 252 |
+
"transcript": "",
|
| 253 |
+
"language": language,
|
| 254 |
+
"model": MODEL_ID,
|
| 255 |
+
"status": "error",
|
| 256 |
+
"error": f"{type(exc).__name__}: {exc}",
|
| 257 |
+
}
|
| 258 |
+
|
| 259 |
+
|
| 260 |
+
def transcribe_text(
|
| 261 |
+
audio_path: str,
|
| 262 |
+
language: str = DEFAULT_LANGUAGE,
|
| 263 |
+
) -> str:
|
| 264 |
+
"""Convenience wrapper that returns just the transcript string.
|
| 265 |
+
|
| 266 |
+
Returns ``""`` on any failure so callers can fall back to typed
|
| 267 |
+
journal input without exception handling.
|
| 268 |
+
"""
|
| 269 |
+
return transcribe(audio_path, language=language).get("transcript", "")
|
app/services/journal.py
CHANGED
|
@@ -62,16 +62,25 @@ LOCATION_KEYWORDS = [
|
|
| 62 |
|
| 63 |
# ββ Transcription βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
|
| 64 |
|
| 65 |
-
def transcribe_journal(
|
|
|
|
|
|
|
|
|
|
|
|
|
| 66 |
"""Transcribe voice journal audio to text.
|
| 67 |
|
| 68 |
-
Tries (in order):
|
| 69 |
-
1. ``
|
| 70 |
-
|
| 71 |
-
|
|
|
|
|
|
|
|
|
|
| 72 |
|
| 73 |
Args:
|
| 74 |
audio_path: Path to a recorded audio file (wav/mp3/m4a/ogg/webm).
|
|
|
|
|
|
|
| 75 |
|
| 76 |
Returns:
|
| 77 |
Transcribed text, or ``""`` if transcription is unavailable.
|
|
@@ -80,7 +89,24 @@ def transcribe_journal(audio_path: str) -> str:
|
|
| 80 |
if not path.exists():
|
| 81 |
raise FileNotFoundError(f"Audio file not found: {audio_path}")
|
| 82 |
|
| 83 |
-
#
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 84 |
try:
|
| 85 |
import whisper # type: ignore
|
| 86 |
|
|
@@ -224,6 +250,9 @@ def create_journal_entry(
|
|
| 224 |
task_id: Optional[str] = None,
|
| 225 |
location_note: str = "",
|
| 226 |
photo_refs: Optional[list[str]] = None,
|
|
|
|
|
|
|
|
|
|
| 227 |
) -> dict:
|
| 228 |
"""Build a complete journal entry dict matching the journal schema.
|
| 229 |
|
|
@@ -234,6 +263,12 @@ def create_journal_entry(
|
|
| 234 |
task_id: Optional associated task ID.
|
| 235 |
location_note: Where the entry was recorded.
|
| 236 |
photo_refs: List of photo identifiers to attach.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 237 |
|
| 238 |
Returns:
|
| 239 |
Full journal entry dict.
|
|
@@ -249,11 +284,18 @@ def create_journal_entry(
|
|
| 249 |
"mood": mood,
|
| 250 |
"location_note": location_note or "Unknown location",
|
| 251 |
"photo_refs": photo_refs or [],
|
|
|
|
| 252 |
}
|
| 253 |
|
| 254 |
if task_id:
|
| 255 |
entry["task_id"] = task_id
|
| 256 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 257 |
return entry
|
| 258 |
|
| 259 |
|
|
|
|
| 62 |
|
| 63 |
# ββ Transcription βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
|
| 64 |
|
| 65 |
+
def transcribe_journal(
|
| 66 |
+
audio_path: str,
|
| 67 |
+
language: str = "en",
|
| 68 |
+
prefer: str = "cohere",
|
| 69 |
+
) -> str:
|
| 70 |
"""Transcribe voice journal audio to text.
|
| 71 |
|
| 72 |
+
Tries (in order, controlled by ``prefer``):
|
| 73 |
+
1. ``"cohere"`` (default) β ``CohereLabs/cohere-transcribe-03-2026``
|
| 74 |
+
via ``app.services.asr``. Lazy-loaded; honors the
|
| 75 |
+
``CITYQUEST_SKIP_MODEL`` / ``CITYQUEST_FAST_TEST`` env vars.
|
| 76 |
+
2. ``"whisper"`` β OpenAI Whisper, runs locally.
|
| 77 |
+
3. Returns an empty string with a warning so the caller can fall
|
| 78 |
+
back to typed input.
|
| 79 |
|
| 80 |
Args:
|
| 81 |
audio_path: Path to a recorded audio file (wav/mp3/m4a/ogg/webm).
|
| 82 |
+
language: Language code passed to the ASR model (e.g. ``"en"``).
|
| 83 |
+
prefer: ``"cohere"`` | ``"whisper"`` | ``"any"``.
|
| 84 |
|
| 85 |
Returns:
|
| 86 |
Transcribed text, or ``""`` if transcription is unavailable.
|
|
|
|
| 89 |
if not path.exists():
|
| 90 |
raise FileNotFoundError(f"Audio file not found: {audio_path}")
|
| 91 |
|
| 92 |
+
# ββ 1. Cohere Transcribe (preferred) ββββββββββββββββββββββββββββββββββ
|
| 93 |
+
if prefer in ("cohere", "any"):
|
| 94 |
+
try:
|
| 95 |
+
from app.services.asr import transcribe_text as _cohere_transcribe
|
| 96 |
+
|
| 97 |
+
text = _cohere_transcribe(str(path), language=language)
|
| 98 |
+
if text:
|
| 99 |
+
return text
|
| 100 |
+
except ImportError as exc:
|
| 101 |
+
print(f"[journal] Cohere ASR import failed: {exc}")
|
| 102 |
+
except Exception as exc:
|
| 103 |
+
print(f"[journal] Cohere ASR failed: {type(exc).__name__}: {exc}")
|
| 104 |
+
|
| 105 |
+
if prefer == "cohere":
|
| 106 |
+
# Don't fall back to Whisper unless the caller asks for it.
|
| 107 |
+
return ""
|
| 108 |
+
|
| 109 |
+
# ββ 2. Whisper fallback (opt-in) ββββββββββββββββββββββββββββββββββββββ
|
| 110 |
try:
|
| 111 |
import whisper # type: ignore
|
| 112 |
|
|
|
|
| 250 |
task_id: Optional[str] = None,
|
| 251 |
location_note: str = "",
|
| 252 |
photo_refs: Optional[list[str]] = None,
|
| 253 |
+
audio_ref: Optional[str] = None,
|
| 254 |
+
asr_metadata: Optional[dict] = None,
|
| 255 |
+
transcript_source: str = "typed",
|
| 256 |
) -> dict:
|
| 257 |
"""Build a complete journal entry dict matching the journal schema.
|
| 258 |
|
|
|
|
| 263 |
task_id: Optional associated task ID.
|
| 264 |
location_note: Where the entry was recorded.
|
| 265 |
photo_refs: List of photo identifiers to attach.
|
| 266 |
+
audio_ref: Optional path/URL of the recorded audio clip.
|
| 267 |
+
asr_metadata: Optional dict describing the ASR pass that
|
| 268 |
+
produced ``transcript`` (model, language, status, error).
|
| 269 |
+
transcript_source: One of ``"typed"`` | ``"asr"`` | ``"hybrid"`` β
|
| 270 |
+
whether the transcript came from typing, the ASR service,
|
| 271 |
+
or a combination (ASR with manual edits).
|
| 272 |
|
| 273 |
Returns:
|
| 274 |
Full journal entry dict.
|
|
|
|
| 284 |
"mood": mood,
|
| 285 |
"location_note": location_note or "Unknown location",
|
| 286 |
"photo_refs": photo_refs or [],
|
| 287 |
+
"transcript_source": transcript_source,
|
| 288 |
}
|
| 289 |
|
| 290 |
if task_id:
|
| 291 |
entry["task_id"] = task_id
|
| 292 |
|
| 293 |
+
if audio_ref:
|
| 294 |
+
entry["audio_ref"] = audio_ref
|
| 295 |
+
|
| 296 |
+
if asr_metadata:
|
| 297 |
+
entry["asr"] = asr_metadata
|
| 298 |
+
|
| 299 |
return entry
|
| 300 |
|
| 301 |
|
requirements.txt
CHANGED
|
@@ -11,4 +11,9 @@ jsonschema
|
|
| 11 |
diffusers[torch] # uncomment for offline FLUX poster generation
|
| 12 |
transformers
|
| 13 |
accelerate
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 14 |
# modal # uncomment to deploy to Modal cloud GPU
|
|
|
|
| 11 |
diffusers[torch] # uncomment for offline FLUX poster generation
|
| 12 |
transformers
|
| 13 |
accelerate
|
| 14 |
+
# Cohere Transcribe ASR (voice journal) β gated model, requires HF_TOKEN
|
| 15 |
+
soundfile
|
| 16 |
+
librosa
|
| 17 |
+
sentencepiece
|
| 18 |
+
protobuf
|
| 19 |
# modal # uncomment to deploy to Modal cloud GPU
|
test_asr.py
ADDED
|
@@ -0,0 +1,249 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Tests for the ASR-backed journal pipeline.
|
| 2 |
+
|
| 3 |
+
These tests cover the fast/skip-model path and the metadata path so the
|
| 4 |
+
rest of the journal pipeline (scoring, recap, persistence) keeps
|
| 5 |
+
working even when no audio model is available.
|
| 6 |
+
|
| 7 |
+
Run with:
|
| 8 |
+
|
| 9 |
+
CITYQUEST_FAST_TEST=1 python test_asr.py
|
| 10 |
+
|
| 11 |
+
To exercise the real model, run without the skip env var on a GPU box.
|
| 12 |
+
"""
|
| 13 |
+
|
| 14 |
+
from __future__ import annotations
|
| 15 |
+
|
| 16 |
+
import json
|
| 17 |
+
import os
|
| 18 |
+
import sys
|
| 19 |
+
import tempfile
|
| 20 |
+
import uuid
|
| 21 |
+
from pathlib import Path
|
| 22 |
+
|
| 23 |
+
|
| 24 |
+
# ββ Fast-test guard βββββββββββββββββββββββββββββββββββββββββββββββββββββββ
|
| 25 |
+
if os.environ.get("CITYQUEST_FAST_TEST") or not os.environ.get("CITYQUEST_RUN_ASR"):
|
| 26 |
+
os.environ.setdefault("CITYQUEST_SKIP_MODEL", "1")
|
| 27 |
+
|
| 28 |
+
passed = 0
|
| 29 |
+
failed = 0
|
| 30 |
+
errors: list[str] = []
|
| 31 |
+
|
| 32 |
+
|
| 33 |
+
def check(label: str, condition: bool, detail: str = "") -> None:
|
| 34 |
+
global passed, failed
|
| 35 |
+
if condition:
|
| 36 |
+
passed += 1
|
| 37 |
+
print(f" β PASS: {label}")
|
| 38 |
+
else:
|
| 39 |
+
failed += 1
|
| 40 |
+
msg = f"β FAIL: {label} β {detail}" if detail else f"β FAIL: {label}"
|
| 41 |
+
errors.append(msg)
|
| 42 |
+
print(f" {msg}")
|
| 43 |
+
|
| 44 |
+
|
| 45 |
+
def main() -> int:
|
| 46 |
+
global passed, failed
|
| 47 |
+
print("=" * 80)
|
| 48 |
+
print("ASR JOURNAL PIPELINE β TESTS")
|
| 49 |
+
print("=" * 80)
|
| 50 |
+
|
| 51 |
+
# ββ T1: ASR module imports & constants ββββββββββββββββββββββββββββββ
|
| 52 |
+
print("\n" + "-" * 80)
|
| 53 |
+
print("T1: ASR MODULE IMPORTS")
|
| 54 |
+
print("-" * 80)
|
| 55 |
+
try:
|
| 56 |
+
from app.services.asr import (
|
| 57 |
+
MODEL_ID,
|
| 58 |
+
SUPPORTED_LANGUAGES,
|
| 59 |
+
DEFAULT_LANGUAGE,
|
| 60 |
+
transcribe,
|
| 61 |
+
transcribe_text,
|
| 62 |
+
)
|
| 63 |
+
check("ASR module imports", True)
|
| 64 |
+
check("Model id is Cohere Transcribe", "cohere-transcribe" in MODEL_ID,
|
| 65 |
+
f"Got {MODEL_ID}")
|
| 66 |
+
check("Supports 14 languages", len(SUPPORTED_LANGUAGES) == 14,
|
| 67 |
+
f"Got {len(SUPPORTED_LANGUAGES)}")
|
| 68 |
+
check("Default language is English", DEFAULT_LANGUAGE == "en")
|
| 69 |
+
except Exception as e:
|
| 70 |
+
check("ASR module imports", False, str(e))
|
| 71 |
+
_summary()
|
| 72 |
+
return 1
|
| 73 |
+
|
| 74 |
+
# ββ T2: transcribe() returns structured result ββββββββββββββββββββββ
|
| 75 |
+
print("\n" + "-" * 80)
|
| 76 |
+
print("T2: transcribe() CONTRACT")
|
| 77 |
+
print("-" * 80)
|
| 78 |
+
try:
|
| 79 |
+
# Use a fake file path β skipped mode should never reach the disk
|
| 80 |
+
result = transcribe("/nonexistent/audio.wav", language="en")
|
| 81 |
+
check("transcribe() returns dict", isinstance(result, dict))
|
| 82 |
+
for key in ("transcript", "language", "model", "status", "error"):
|
| 83 |
+
check(f"transcribe() has '{key}'", key in result)
|
| 84 |
+
check("status is 'ok' | 'skipped' | 'error'",
|
| 85 |
+
result["status"] in ("ok", "skipped", "error"),
|
| 86 |
+
f"Got {result['status']}")
|
| 87 |
+
check("transcript is a string", isinstance(result["transcript"], str))
|
| 88 |
+
check("model is reported", bool(result["model"]))
|
| 89 |
+
except Exception as e:
|
| 90 |
+
check("transcribe() contract", False, str(e))
|
| 91 |
+
|
| 92 |
+
# ββ T3: Skipped mode is silent and empty ββββββββββββββββββββββββββββ
|
| 93 |
+
print("\n" + "-" * 80)
|
| 94 |
+
print("T3: SKIP-MODE BEHAVIOR")
|
| 95 |
+
print("-" * 80)
|
| 96 |
+
try:
|
| 97 |
+
os.environ["CITYQUEST_SKIP_MODEL"] = "1"
|
| 98 |
+
result = transcribe("/tmp/fake.wav", language="en")
|
| 99 |
+
check("Skip mode returns empty transcript", result["transcript"] == "")
|
| 100 |
+
check("Skip mode reports status=skipped", result["status"] == "skipped")
|
| 101 |
+
# Restore to be safe
|
| 102 |
+
del os.environ["CITYQUEST_SKIP_MODEL"]
|
| 103 |
+
except Exception as e:
|
| 104 |
+
check("Skip-mode behavior", False, str(e))
|
| 105 |
+
|
| 106 |
+
# ββ T4: transcribe_journal() returns empty on skip ββββββββββββββββββ
|
| 107 |
+
print("\n" + "-" * 80)
|
| 108 |
+
print("T4: JOURNAL.transcribe_journal() FALLBACK")
|
| 109 |
+
print("-" * 80)
|
| 110 |
+
try:
|
| 111 |
+
os.environ["CITYQUEST_SKIP_MODEL"] = "1"
|
| 112 |
+
from app.services.journal import transcribe_journal
|
| 113 |
+
|
| 114 |
+
# Create a tiny valid wav so path checks pass
|
| 115 |
+
with tempfile.NamedTemporaryFile(suffix=".wav", delete=False) as fh:
|
| 116 |
+
wav_path = fh.name
|
| 117 |
+
# Minimal RIFF header β not a real wav, but enough for the
|
| 118 |
+
# existence check in transcribe_journal()
|
| 119 |
+
fh.write(b"RIFF\x24\x00\x00\x00WAVEfmt ")
|
| 120 |
+
|
| 121 |
+
try:
|
| 122 |
+
text = transcribe_journal(wav_path, language="en")
|
| 123 |
+
check("transcribe_journal returns str", isinstance(text, str))
|
| 124 |
+
check("transcribe_journal empty when ASR skipped", text == "",
|
| 125 |
+
f"Got {text!r}")
|
| 126 |
+
finally:
|
| 127 |
+
try:
|
| 128 |
+
os.unlink(wav_path)
|
| 129 |
+
except OSError:
|
| 130 |
+
pass
|
| 131 |
+
except Exception as e:
|
| 132 |
+
check("Journal.transcribe_journal fallback", False, str(e))
|
| 133 |
+
|
| 134 |
+
# ββ T5: create_journal_entry() supports ASR metadata βββββββββββββββ
|
| 135 |
+
print("\n" + "-" * 80)
|
| 136 |
+
print("T5: JOURNAL ENTRY METADATA")
|
| 137 |
+
print("-" * 80)
|
| 138 |
+
try:
|
| 139 |
+
from app.services.journal import create_journal_entry, save_journal_entry, load_journal_entries
|
| 140 |
+
|
| 141 |
+
asr_meta = {
|
| 142 |
+
"model": "CohereLabs/cohere-transcribe-03-2026",
|
| 143 |
+
"language": "en",
|
| 144 |
+
"status": "ok",
|
| 145 |
+
"error": None,
|
| 146 |
+
}
|
| 147 |
+
entry = create_journal_entry(
|
| 148 |
+
transcript="We just found the mural near the canal, amazing!",
|
| 149 |
+
session_id="test-asr",
|
| 150 |
+
team_id="team-a",
|
| 151 |
+
task_id="t1",
|
| 152 |
+
location_note="Canal area",
|
| 153 |
+
audio_ref="/tmp/clip.wav",
|
| 154 |
+
asr_metadata=asr_meta,
|
| 155 |
+
transcript_source="asr",
|
| 156 |
+
)
|
| 157 |
+
check("Entry has transcript_source", entry.get("transcript_source") == "asr")
|
| 158 |
+
check("Entry has audio_ref", entry.get("audio_ref") == "/tmp/clip.wav")
|
| 159 |
+
check("Entry has asr metadata", isinstance(entry.get("asr"), dict))
|
| 160 |
+
check("ASR metadata has model", entry["asr"]["model"] == asr_meta["model"])
|
| 161 |
+
check("ASR metadata has language", entry["asr"]["language"] == "en")
|
| 162 |
+
check("ASR metadata has status", entry["asr"]["status"] == "ok")
|
| 163 |
+
|
| 164 |
+
# Schema validation (optional β only if jsonschema is installed)
|
| 165 |
+
try:
|
| 166 |
+
import jsonschema # type: ignore
|
| 167 |
+
schema_path = Path("app/schemas/journal_schema.json")
|
| 168 |
+
schema = json.loads(schema_path.read_text())
|
| 169 |
+
jsonschema.validate(instance=entry, schema=schema)
|
| 170 |
+
check("Entry validates against journal_schema.json", True)
|
| 171 |
+
except ImportError:
|
| 172 |
+
check("Entry validates against journal_schema.json (skipped, no jsonschema)", True)
|
| 173 |
+
|
| 174 |
+
# Persistence round-trip
|
| 175 |
+
save_journal_entry(entry)
|
| 176 |
+
loaded = load_journal_entries(session_id="test-asr")
|
| 177 |
+
check("Entry persisted with ASR metadata", len(loaded) >= 1)
|
| 178 |
+
if loaded:
|
| 179 |
+
check("Loaded entry has transcript_source",
|
| 180 |
+
loaded[-1].get("transcript_source") == "asr")
|
| 181 |
+
check("Loaded entry has asr block",
|
| 182 |
+
isinstance(loaded[-1].get("asr"), dict))
|
| 183 |
+
except Exception as e:
|
| 184 |
+
check("Journal entry metadata", False, str(e))
|
| 185 |
+
|
| 186 |
+
# ββ T6: app.record_journal() handles voice path βββββββββββββββββββββ
|
| 187 |
+
print("\n" + "-" * 80)
|
| 188 |
+
print("T6: app.record_journal() VOICE PATH (skip-mode)")
|
| 189 |
+
print("-" * 80)
|
| 190 |
+
try:
|
| 191 |
+
os.environ["CITYQUEST_SKIP_MODEL"] = "1"
|
| 192 |
+
# We can't import the whole app (it boots Gradio), so test the
|
| 193 |
+
# logic by directly exercising the journal functions.
|
| 194 |
+
from app.services.journal import transcribe_journal, create_journal_entry
|
| 195 |
+
from app.services.asr import transcribe as _asr_transcribe
|
| 196 |
+
|
| 197 |
+
with tempfile.NamedTemporaryFile(suffix=".wav", delete=False) as fh:
|
| 198 |
+
wav_path = fh.name
|
| 199 |
+
fh.write(b"RIFF\x24\x00\x00\x00WAVEfmt ")
|
| 200 |
+
|
| 201 |
+
try:
|
| 202 |
+
asr_result = _asr_transcribe(wav_path, language="en")
|
| 203 |
+
check("ASR returns skipped status", asr_result["status"] == "skipped")
|
| 204 |
+
check("ASR transcript is empty", asr_result["transcript"] == "")
|
| 205 |
+
|
| 206 |
+
# When ASR produces no text, journal creation should still
|
| 207 |
+
# be possible via the typed/manual correction path.
|
| 208 |
+
entry = create_journal_entry(
|
| 209 |
+
transcript="Typed correction (ASR was unavailable).",
|
| 210 |
+
session_id="test-asr-voice",
|
| 211 |
+
audio_ref=wav_path,
|
| 212 |
+
asr_metadata={
|
| 213 |
+
"model": asr_result["model"],
|
| 214 |
+
"language": asr_result["language"],
|
| 215 |
+
"status": asr_result["status"],
|
| 216 |
+
"error": asr_result["error"],
|
| 217 |
+
},
|
| 218 |
+
transcript_source="typed",
|
| 219 |
+
)
|
| 220 |
+
check("Hybrid entry created with asr metadata",
|
| 221 |
+
entry.get("asr", {}).get("status") == "skipped")
|
| 222 |
+
finally:
|
| 223 |
+
try:
|
| 224 |
+
os.unlink(wav_path)
|
| 225 |
+
except OSError:
|
| 226 |
+
pass
|
| 227 |
+
except Exception as e:
|
| 228 |
+
check("Voice path in skip-mode", False, str(e))
|
| 229 |
+
|
| 230 |
+
_summary()
|
| 231 |
+
return 0 if failed == 0 else 1
|
| 232 |
+
|
| 233 |
+
|
| 234 |
+
def _summary() -> None:
|
| 235 |
+
total = passed + failed
|
| 236 |
+
print("\n" + "=" * 80)
|
| 237 |
+
if failed == 0:
|
| 238 |
+
print(f"RESULTS: {passed}/{total} tests passed β ALL CLEAR π")
|
| 239 |
+
else:
|
| 240 |
+
print(f"RESULTS: {passed}/{total} tests passed β {failed} FAILED")
|
| 241 |
+
print("=" * 80)
|
| 242 |
+
if errors:
|
| 243 |
+
print("\nFailed tests:")
|
| 244 |
+
for e in errors:
|
| 245 |
+
print(f" {e}")
|
| 246 |
+
|
| 247 |
+
|
| 248 |
+
if __name__ == "__main__":
|
| 249 |
+
sys.exit(main())
|