NANI-Nithin commited on
Commit
e7db016
Β·
1 Parent(s): 4f6f969

feat: Added voice journal recording with Cohere ASR integration

Browse files
COHERE_ASR_SETUP.md ADDED
@@ -0,0 +1,127 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Cohere Transcribe ASR β€” Setup Guide
2
+
3
+ ## Overview
4
+
5
+ The voice journal pipeline uses **CohereLabs/cohere-transcribe-03-2026**, a
6
+ 2 B-parameter conformer encoder + lightweight transformer decoder trained
7
+ from scratch for ASR. It is gated on the Hugging Face Hub (you must
8
+ accept the model terms once with your account) and supports 14
9
+ languages:
10
+
11
+ - **European:** English, French, German, Italian, Spanish, Portuguese,
12
+ Greek, Dutch, Polish
13
+ - **APAC:** Chinese (Mandarin), Japanese, Korean, Vietnamese
14
+ - **MENA:** Arabic
15
+
16
+ The model is Apache 2.0 licensed and is integrated into the journal
17
+ pipeline so players can speak during a game instead of typing.
18
+
19
+ ## Why this configuration
20
+
21
+ 1. **Sponsor visibility** β€” Cohere Labs is a hackathon sponsor.
22
+ 2. **State-of-the-art accuracy** β€” 5.42 mean WER on the Open ASR
23
+ Leaderboard (5.x–10.x WER across real-world domains) and 1.25 WER on
24
+ LibriSpeech clean.
25
+ 3. **Production runtime** β€” supports πŸ€— Transformers (offline),
26
+ vLLM, mlx-audio, Rust, and a WebGPU browser demo.
27
+ 4. **Lazy loading** β€” the model is downloaded on first use, never at
28
+ app startup, so demo boot is unaffected.
29
+
30
+ ## Installation
31
+
32
+ ### 1. Accept the model terms
33
+
34
+ Visit <https://huggingface.co/CohereLabs/cohere-transcribe-03-2026>,
35
+ click **Agree and access repository** with the account you plan to
36
+ authenticate as.
37
+
38
+ ### 2. Install dependencies
39
+
40
+ ```bash
41
+ pip install 'transformers>=5.4.0' torch huggingface_hub \
42
+ soundfile librosa sentencepiece protobuf
43
+ ```
44
+
45
+ (These are added to `requirements.txt` for the demo; `transformers` is
46
+ already present.)
47
+
48
+ ### 3. Provide an HF token
49
+
50
+ Set a token in the environment so the gated model can be downloaded:
51
+
52
+ ```bash
53
+ export HF_TOKEN=hf_xxx... # Linux/macOS
54
+ $env:HF_TOKEN="hf_xxx..." # PowerShell
55
+ ```
56
+
57
+ On Hugging Face Spaces, create a `HF_TOKEN` secret (same name as
58
+ `huggingface` in `modal_serve.py`/`modal_train.py`).
59
+
60
+ ## How the pipeline uses it
61
+
62
+ ```text
63
+ Gradio microphone / upload
64
+ ↓
65
+ app.py: record_journal(audio_path, language)
66
+ ↓
67
+ app/services/asr.py: transcribe(audio_path, language)
68
+ ↓
69
+ CohereAsrForConditionalGeneration ← CohereLabs/cohere-transcribe-03-2026
70
+ ↓
71
+ transcript
72
+ ↓
73
+ app/services/journal.py: create_journal_entry(...)
74
+ ↓
75
+ app/logs/journals.jsonl
76
+ ```
77
+
78
+ Each journal entry now carries:
79
+
80
+ - `transcript_source` β€” `"typed" | "asr" | "hybrid"`
81
+ - `audio_ref` β€” path of the recorded audio clip
82
+ - `asr` β€” `{ model, language, status, error }`
83
+
84
+ The `journal_recorded` event log also includes `transcript_source`,
85
+ `asr_status`, and `asr_model` for full traceability.
86
+
87
+ ## Skipping the model in tests
88
+
89
+ To run the demo without downloading the model, set either of:
90
+
91
+ ```bash
92
+ CITYQUEST_SKIP_MODEL=1
93
+ CITYQUEST_FAST_TEST=1
94
+ ```
95
+
96
+ When set, `app.services.asr.transcribe()` returns
97
+ `status="skipped"` with an empty transcript. The journal pipeline
98
+ silently falls back to typed input.
99
+
100
+ ## Verification
101
+
102
+ ```bash
103
+ $env:CITYQUEST_FAST_TEST="1"
104
+ .\.venv\Scripts\python.exe test_asr.py
105
+ .\.venv\Scripts\python.exe test_end_to_end.py
106
+ ```
107
+
108
+ Expected: 30/30 ASR tests + 86/86 end-to-end tests pass in skip-mode.
109
+
110
+ ## Limitations (per the model card)
111
+
112
+ 1. **Single language per call** β€” pick the right language code; the
113
+ model does not auto-detect or handle code-switching well.
114
+ 2. **No diarization or timestamps** β€” only plain text is returned.
115
+ 3. **Eager on silence** β€” prepend a VAD/silence gate if the recording
116
+ has noisy backgrounds; otherwise the model may hallucinate.
117
+
118
+ ## File map
119
+
120
+ | File | Purpose |
121
+ | --- | --- |
122
+ | `app/services/asr.py` | Lazy-loaded Cohere Transcribe wrapper. |
123
+ | `app/services/journal.py` | `transcribe_journal()` and `create_journal_entry()` now accept ASR metadata. |
124
+ | `app/schemas/journal_schema.json` | Optional `transcript_source`, `audio_ref`, `asr` fields. |
125
+ | `app.py` | Gradio audio component + `record_journal()` voice path. |
126
+ | `test_asr.py` | Skip-mode tests for the ASR pipeline. |
127
+ | `requirements.txt` | Optional ASR runtime deps. |
README.md CHANGED
@@ -28,7 +28,20 @@ An AI-powered platform that generates complete, playable real-world games β€” ru
28
  1. Select a game type, city, and preferences
29
  2. AI generates rules, tasks, and hints
30
  3. Play in the real world, track progress in the app
31
- 4. Receive an AI-generated story summary of the outcome
 
 
 
 
 
 
 
 
 
 
 
 
 
32
 
33
  ## Team
34
 
 
28
  1. Select a game type, city, and preferences
29
  2. AI generates rules, tasks, and hints
30
  3. Play in the real world, track progress in the app
31
+ 4. Record voice journals during play (auto-transcribed with [Cohere Transcribe](https://huggingface.co/CohereLabs/cohere-transcribe-03-2026))
32
+ 5. Receive an AI-generated story summary of the outcome
33
+
34
+ ## AI Models
35
+
36
+ | Stage | Model | Notes |
37
+ | --- | --- | --- |
38
+ | Game generation | `nvidia/NVIDIA-Nemotron-3-Nano-4B-GGUF` via llama.cpp | Retrieval-grounded, validation + repair enforced |
39
+ | Voice journal ASR | `CohereLabs/cohere-transcribe-03-2026` via πŸ€— Transformers | 14 languages, lazy-loaded; typed-input fallback |
40
+ | Recap (optional) | `openbmb/MiniCPM5-1B-GGUF` | Currently using deterministic template recap for reliability |
41
+ | Poster (optional) | `black-forest-labs/FLUX.1-schnell` | Skipped under `CITYQUEST_SKIP_MODEL` |
42
+
43
+ Set `CITYQUEST_FAST_TEST=1` to run the entire pipeline without
44
+ downloading any model weights.
45
 
46
  ## Team
47
 
app.py CHANGED
@@ -215,22 +215,77 @@ def use_hint(session_id: str, task_id: str, team_id: str = "team-a"):
215
 
216
  def record_journal(
217
  session_id: str,
218
- transcript: str,
219
  task_id: str = "",
220
  location_note: str = "",
221
  team_id: str = "team-a",
 
 
222
  ):
223
- """Record a text journal entry, summarize it, and return the result."""
 
 
 
 
 
 
 
 
 
 
 
 
 
224
  if session_id not in SESSION_STORE:
225
  return "[!] Unknown session"
226
 
227
- # Build full journal entry
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
228
  entry = create_journal_entry(
229
  transcript=transcript,
230
  session_id=session_id,
231
  team_id=team_id,
232
  task_id=task_id or None,
233
  location_note=location_note,
 
 
 
234
  )
235
 
236
  # Summarize
@@ -248,16 +303,28 @@ def record_journal(
248
  "mood": entry["mood"],
249
  "story_value": summary["story_value"],
250
  "summary": summary["moment_summary"],
 
 
 
251
  }, team_id=team_id)
252
  SESSION_STORE[session_id]["events"].append(ev)
253
  SESSION_STORE[session_id]["journals"].append(entry)
254
 
255
  # Build display text
 
 
 
 
 
 
 
 
256
  display = (
257
- f"πŸŽ™οΈ **Journal recorded!**\n"
258
  f"- Mood: *{entry['mood']}*\n"
259
  f"- Story value: **{summary['story_value']}**\n"
260
  f"- Tags: {', '.join(summary['tags'])}\n"
 
261
  f"- Summary: {summary['moment_summary']}"
262
  )
263
  return display
@@ -1008,13 +1075,24 @@ with gr.Blocks(title="CityQuest-AI") as demo:
1008
  task_feedback = gr.Markdown(value="")
1009
 
1010
  gr.HTML('<div class="section-label" style="margin-top:20px;">Voice Journal</div>')
 
 
 
 
 
1011
  journal_transcript = gr.Textbox(
1012
- label="What happened? How do you feel?", lines=3,
 
1013
  placeholder="We just found the mural near the canal β€” it was incredible!",
1014
  )
1015
  with gr.Row():
1016
  journal_task_id = gr.Textbox(label="Task ID (opt.)", placeholder="t1")
1017
  journal_location = gr.Textbox(label="Location note", placeholder="Rue de Rivoli")
 
 
 
 
 
1018
  journal_btn = gr.Button("πŸŽ™οΈ Record Journal", variant="secondary")
1019
  journal_output = gr.Markdown(value="")
1020
 
@@ -1226,7 +1304,8 @@ with gr.Blocks(title="CityQuest-AI") as demo:
1226
 
1227
  journal_btn.click(
1228
  fn=record_journal,
1229
- inputs=[current_session, journal_transcript, journal_task_id, journal_location, current_team],
 
1230
  outputs=[journal_output],
1231
  )
1232
  photo_btn.click(
 
215
 
216
  def record_journal(
217
  session_id: str,
218
+ transcript: str = "",
219
  task_id: str = "",
220
  location_note: str = "",
221
  team_id: str = "team-a",
222
+ audio_path: str = "",
223
+ language: str = "en",
224
  ):
225
+ """Record a journal entry, summarize it, and return the result.
226
+
227
+ Two input paths are supported:
228
+
229
+ * **Voice path** β€” pass ``audio_path`` (e.g. a path returned by
230
+ ``gr.Audio(type="filepath")``). The audio is transcribed with the
231
+ Cohere ASR service (``CohereLabs/cohere-transcribe-03-2026``) and
232
+ the transcript is stored with ASR metadata.
233
+ * **Typed path** β€” pass ``transcript`` directly. The user can also
234
+ edit an ASR transcript in the UI before submitting; if both
235
+ ``audio_path`` and ``transcript`` are present, the typed
236
+ transcript wins and is treated as a manual correction
237
+ (``transcript_source == "hybrid"``).
238
+ """
239
  if session_id not in SESSION_STORE:
240
  return "[!] Unknown session"
241
 
242
+ asr_metadata: dict | None = None
243
+ audio_ref: str | None = None
244
+ transcript_source = "typed"
245
+
246
+ # ── 1. Voice path β€” transcribe audio if provided ───────────────────
247
+ if audio_path:
248
+ audio_ref = audio_path
249
+ try:
250
+ from app.services.asr import transcribe as _asr_transcribe
251
+
252
+ asr_result = _asr_transcribe(audio_path, language=language)
253
+ asr_metadata = {
254
+ "model": asr_result.get("model"),
255
+ "language": asr_result.get("language"),
256
+ "status": asr_result.get("status"),
257
+ "error": asr_result.get("error"),
258
+ }
259
+ asr_text = (asr_result.get("transcript") or "").strip()
260
+ if asr_text and not transcript.strip():
261
+ transcript = asr_text
262
+ transcript_source = "asr"
263
+ elif asr_text and transcript.strip() and asr_text != transcript.strip():
264
+ transcript_source = "hybrid"
265
+ except Exception as exc:
266
+ print(f"[app] ASR transcription failed: {type(exc).__name__}: {exc}")
267
+ asr_metadata = {
268
+ "model": None,
269
+ "language": language,
270
+ "status": "error",
271
+ "error": f"{type(exc).__name__}: {exc}",
272
+ }
273
+
274
+ if not transcript or not transcript.strip():
275
+ return (
276
+ "⚠️ No transcript available β€” record audio or type a note first."
277
+ )
278
+
279
+ # ── 2. Build journal entry ────────────────────────────────────────
280
  entry = create_journal_entry(
281
  transcript=transcript,
282
  session_id=session_id,
283
  team_id=team_id,
284
  task_id=task_id or None,
285
  location_note=location_note,
286
+ audio_ref=audio_ref,
287
+ asr_metadata=asr_metadata,
288
+ transcript_source=transcript_source,
289
  )
290
 
291
  # Summarize
 
303
  "mood": entry["mood"],
304
  "story_value": summary["story_value"],
305
  "summary": summary["moment_summary"],
306
+ "transcript_source": transcript_source,
307
+ "asr_status": asr_metadata.get("status") if asr_metadata else None,
308
+ "asr_model": asr_metadata.get("model") if asr_metadata else None,
309
  }, team_id=team_id)
310
  SESSION_STORE[session_id]["events"].append(ev)
311
  SESSION_STORE[session_id]["journals"].append(entry)
312
 
313
  # Build display text
314
+ source_label = {
315
+ "asr": "πŸŽ™οΈ (transcribed)",
316
+ "hybrid": "πŸŽ™οΈβœοΈ (transcribed, edited)",
317
+ "typed": "⌨️ (typed)",
318
+ }.get(transcript_source, transcript_source)
319
+ asr_line = ""
320
+ if asr_metadata and asr_metadata.get("status") != "ok":
321
+ asr_line = f"- ASR status: **{asr_metadata.get('status')}** β€” {asr_metadata.get('error') or ''}\n"
322
  display = (
323
+ f"πŸŽ™οΈ **Journal recorded!** {source_label}\n"
324
  f"- Mood: *{entry['mood']}*\n"
325
  f"- Story value: **{summary['story_value']}**\n"
326
  f"- Tags: {', '.join(summary['tags'])}\n"
327
+ f"{asr_line}"
328
  f"- Summary: {summary['moment_summary']}"
329
  )
330
  return display
 
1075
  task_feedback = gr.Markdown(value="")
1076
 
1077
  gr.HTML('<div class="section-label" style="margin-top:20px;">Voice Journal</div>')
1078
+ journal_audio = gr.Audio(
1079
+ label="πŸŽ™οΈ Record voice (auto-transcribed via Cohere ASR)",
1080
+ sources=["microphone", "upload"],
1081
+ type="filepath",
1082
+ )
1083
  journal_transcript = gr.Textbox(
1084
+ label="Transcript (edit the ASR result or type your own)",
1085
+ lines=3,
1086
  placeholder="We just found the mural near the canal β€” it was incredible!",
1087
  )
1088
  with gr.Row():
1089
  journal_task_id = gr.Textbox(label="Task ID (opt.)", placeholder="t1")
1090
  journal_location = gr.Textbox(label="Location note", placeholder="Rue de Rivoli")
1091
+ journal_language = gr.Dropdown(
1092
+ label="Language", value="en",
1093
+ choices=["en", "fr", "de", "it", "es", "pt", "el", "nl", "pl",
1094
+ "zh", "ja", "ko", "vi", "ar"],
1095
+ )
1096
  journal_btn = gr.Button("πŸŽ™οΈ Record Journal", variant="secondary")
1097
  journal_output = gr.Markdown(value="")
1098
 
 
1304
 
1305
  journal_btn.click(
1306
  fn=record_journal,
1307
+ inputs=[current_session, journal_transcript, journal_task_id, journal_location,
1308
+ current_team, journal_audio, journal_language],
1309
  outputs=[journal_output],
1310
  )
1311
  photo_btn.click(
app/data/sessions_store.json CHANGED
@@ -1 +1 @@
1
- {"codes": {"75ZKFE": "9fcb0953-4e36-4485-9bbe-fad577de31fa", "JWR5RV": "6b1af473-b00b-44ae-bffc-7b8e93616767"}, "sessions": {"9fcb0953-4e36-4485-9bbe-fad577de31fa": {"config": {"game_type": "scavenger_hunt", "city": "Paris", "area": "Le Marais", "location_type": "mixed", "duration_minutes": 60, "num_players": 2, "difficulty": "easy", "age_group": "adults", "energy_level": "low", "photo_enabled": true}, "game": {"game_id": "mock-8506e14a", "title": "Scavenger_Hunt in Le Marais", "theme": "easy adventure", "setup": {"city": "Paris", "area": "Le Marais", "meeting_point": "Main entrance of Le Marais", "duration_minutes": 60, "num_players": 2}, "rules": ["Complete as many tasks as possible within 60 minutes", "Take photos or notes as proof of completion", "Stay within the designated area at all times", "No entering private buildings or restricted areas", "This game is suitable for adults"], "tasks": [{"task_id": "t1", "title": "Task 1: Explore the main square", "description": "Find and document something interesting in the main square", "location_hint": "Navigate to the main square and look for distinctive features", "points": 15, "time_limit_minutes": 8, "proof_type": "photo", "hint": "Look for signs or landmarks in the main square", "safety_note": "Stay on public paths and avoid restricted areas"}, {"task_id": "t2", "title": "Task 2: Explore the city center", "description": "Find and document something interesting in the city center", "location_hint": "Navigate to the city center and look for distinctive features", "points": 20, "time_limit_minutes": 10, "proof_type": "observation", "hint": "Look for signs or landmarks in the city center", "safety_note": "Stay on public paths and avoid restricted areas"}, {"task_id": "t3", "title": "Task 3: Explore the park area", "description": "Find and document something interesting in the park area", "location_hint": "Navigate to the park area and look for distinctive features", "points": 25, "time_limit_minutes": 12, "proof_type": "text", "hint": "Look for signs or landmarks in the park area", "safety_note": "Stay on public paths and avoid restricted areas"}, {"task_id": "t4", "title": "Task 4: Explore the landmark district", "description": "Find and document something interesting in the landmark district", "location_hint": "Navigate to the landmark district and look for distinctive features", "points": 30, "time_limit_minutes": 14, "proof_type": "photo", "hint": "Look for signs or landmarks in the landmark district", "safety_note": "Stay on public paths and avoid restricted areas"}], "global_hints": ["Explore systematically from the meeting point outward", "Ask locals for directions if needed", "Time management is key - don't spend too long on any single task"], "score_rules": ["Each task completed: full points", "Early completion: +1 bonus point per minute under limit", "Hints used: -5 points per hint", "Late arrival at meeting point: -10 points per minute"], "tie_breaker": "Winner is the player with the most points when time expires. Ties broken by earliest completion time.", "safety": {"allowed_zone": "Le Marais", "forbidden_behaviors": ["Entering buildings without permission", "Crossing busy streets recklessly", "Approaching strangers", "Leaving the designated area"], "adult_supervision": false, "stop_conditions": ["If a player feels unsafe, the game stops immediately", "If weather becomes severe, relocate to shelter", "If anyone is injured, call emergency services"]}, "story_seed": {"tone": "playful", "motifs": ["exploration", "discovery", "teamwork"], "recap_style": "episode_recap"}}, "adventure_code": "75ZKFE", "teams": ["team-a", "team-b"], "players": [{"name": "Player 1", "team_id": "team-a"}, {"name": "play 2", "team_id": "team-b"}, {"name": "hola", "team_id": "team-b"}], "photos": [{"photo_id": "photo-71c44ee8", "photo_name": "", "caption": "", "task_id": "t1"}]}, "6b1af473-b00b-44ae-bffc-7b8e93616767": {"config": {"game_type": "scavenger_hunt", "city": "Paris", "area": "Le Marais", "location_type": "mixed", "duration_minutes": 60, "num_players": 2, "difficulty": "easy", "age_group": "adults", "energy_level": "low", "photo_enabled": true}, "game": {"game_id": "mock-6d6ab9cf", "title": "Scavenger_Hunt in Le Marais", "theme": "easy adventure", "setup": {"city": "Paris", "area": "Le Marais", "meeting_point": "Main entrance of Le Marais", "duration_minutes": 60, "num_players": 2}, "rules": ["Complete as many tasks as possible within 60 minutes", "Take photos or notes as proof of completion", "Stay within the designated area at all times", "No entering private buildings or restricted areas", "This game is suitable for adults"], "tasks": [{"task_id": "t1", "title": "Task 1: Explore the main square", "description": "Find and document something interesting in the main square", "location_hint": "Navigate to the main square and look for distinctive features", "points": 15, "time_limit_minutes": 8, "proof_type": "photo", "hint": "Look for signs or landmarks in the main square", "safety_note": "Stay on public paths and avoid restricted areas"}, {"task_id": "t2", "title": "Task 2: Explore the city center", "description": "Find and document something interesting in the city center", "location_hint": "Navigate to the city center and look for distinctive features", "points": 20, "time_limit_minutes": 10, "proof_type": "observation", "hint": "Look for signs or landmarks in the city center", "safety_note": "Stay on public paths and avoid restricted areas"}, {"task_id": "t3", "title": "Task 3: Explore the park area", "description": "Find and document something interesting in the park area", "location_hint": "Navigate to the park area and look for distinctive features", "points": 25, "time_limit_minutes": 12, "proof_type": "text", "hint": "Look for signs or landmarks in the park area", "safety_note": "Stay on public paths and avoid restricted areas"}, {"task_id": "t4", "title": "Task 4: Explore the landmark district", "description": "Find and document something interesting in the landmark district", "location_hint": "Navigate to the landmark district and look for distinctive features", "points": 30, "time_limit_minutes": 14, "proof_type": "photo", "hint": "Look for signs or landmarks in the landmark district", "safety_note": "Stay on public paths and avoid restricted areas"}], "global_hints": ["Explore systematically from the meeting point outward", "Ask locals for directions if needed", "Time management is key - don't spend too long on any single task"], "score_rules": ["Each task completed: full points", "Early completion: +1 bonus point per minute under limit", "Hints used: -5 points per hint", "Late arrival at meeting point: -10 points per minute"], "tie_breaker": "Winner is the player with the most points when time expires. Ties broken by earliest completion time.", "safety": {"allowed_zone": "Le Marais", "forbidden_behaviors": ["Entering buildings without permission", "Crossing busy streets recklessly", "Approaching strangers", "Leaving the designated area"], "adult_supervision": false, "stop_conditions": ["If a player feels unsafe, the game stops immediately", "If weather becomes severe, relocate to shelter", "If anyone is injured, call emergency services"]}, "story_seed": {"tone": "playful", "motifs": ["exploration", "discovery", "teamwork"], "recap_style": "episode_recap"}}, "adventure_code": "JWR5RV", "teams": ["team-a", "team-b"], "players": [{"name": "Player 1", "team_id": "team-a"}], "photos": []}}}
 
1
+ {"codes": {"75ZKFE": "9fcb0953-4e36-4485-9bbe-fad577de31fa", "JWR5RV": "6b1af473-b00b-44ae-bffc-7b8e93616767", "NKPZKT": "32e1bbe3-50a1-4e43-a7c2-8ecf3f781550", "NXWHQK": "4b03fbfc-4326-4325-8975-e4216c44e03e"}, "sessions": {"9fcb0953-4e36-4485-9bbe-fad577de31fa": {"config": {"game_type": "scavenger_hunt", "city": "Paris", "area": "Le Marais", "location_type": "mixed", "duration_minutes": 60, "num_players": 2, "difficulty": "easy", "age_group": "adults", "energy_level": "low", "photo_enabled": true}, "game": {"game_id": "mock-8506e14a", "title": "Scavenger_Hunt in Le Marais", "theme": "easy adventure", "setup": {"city": "Paris", "area": "Le Marais", "meeting_point": "Main entrance of Le Marais", "duration_minutes": 60, "num_players": 2}, "rules": ["Complete as many tasks as possible within 60 minutes", "Take photos or notes as proof of completion", "Stay within the designated area at all times", "No entering private buildings or restricted areas", "This game is suitable for adults"], "tasks": [{"task_id": "t1", "title": "Task 1: Explore the main square", "description": "Find and document something interesting in the main square", "location_hint": "Navigate to the main square and look for distinctive features", "points": 15, "time_limit_minutes": 8, "proof_type": "photo", "hint": "Look for signs or landmarks in the main square", "safety_note": "Stay on public paths and avoid restricted areas"}, {"task_id": "t2", "title": "Task 2: Explore the city center", "description": "Find and document something interesting in the city center", "location_hint": "Navigate to the city center and look for distinctive features", "points": 20, "time_limit_minutes": 10, "proof_type": "observation", "hint": "Look for signs or landmarks in the city center", "safety_note": "Stay on public paths and avoid restricted areas"}, {"task_id": "t3", "title": "Task 3: Explore the park area", "description": "Find and document something interesting in the park area", "location_hint": "Navigate to the park area and look for distinctive features", "points": 25, "time_limit_minutes": 12, "proof_type": "text", "hint": "Look for signs or landmarks in the park area", "safety_note": "Stay on public paths and avoid restricted areas"}, {"task_id": "t4", "title": "Task 4: Explore the landmark district", "description": "Find and document something interesting in the landmark district", "location_hint": "Navigate to the landmark district and look for distinctive features", "points": 30, "time_limit_minutes": 14, "proof_type": "photo", "hint": "Look for signs or landmarks in the landmark district", "safety_note": "Stay on public paths and avoid restricted areas"}], "global_hints": ["Explore systematically from the meeting point outward", "Ask locals for directions if needed", "Time management is key - don't spend too long on any single task"], "score_rules": ["Each task completed: full points", "Early completion: +1 bonus point per minute under limit", "Hints used: -5 points per hint", "Late arrival at meeting point: -10 points per minute"], "tie_breaker": "Winner is the player with the most points when time expires. Ties broken by earliest completion time.", "safety": {"allowed_zone": "Le Marais", "forbidden_behaviors": ["Entering buildings without permission", "Crossing busy streets recklessly", "Approaching strangers", "Leaving the designated area"], "adult_supervision": false, "stop_conditions": ["If a player feels unsafe, the game stops immediately", "If weather becomes severe, relocate to shelter", "If anyone is injured, call emergency services"]}, "story_seed": {"tone": "playful", "motifs": ["exploration", "discovery", "teamwork"], "recap_style": "episode_recap"}}, "adventure_code": "75ZKFE", "teams": ["team-a", "team-b"], "players": [{"name": "Player 1", "team_id": "team-a"}, {"name": "play 2", "team_id": "team-b"}, {"name": "hola", "team_id": "team-b"}], "photos": [{"photo_id": "photo-71c44ee8", "photo_name": "", "caption": "", "task_id": "t1"}]}, "6b1af473-b00b-44ae-bffc-7b8e93616767": {"config": {"game_type": "scavenger_hunt", "city": "Paris", "area": "Le Marais", "location_type": "mixed", "duration_minutes": 60, "num_players": 2, "difficulty": "easy", "age_group": "adults", "energy_level": "low", "photo_enabled": true}, "game": {"game_id": "mock-6d6ab9cf", "title": "Scavenger_Hunt in Le Marais", "theme": "easy adventure", "setup": {"city": "Paris", "area": "Le Marais", "meeting_point": "Main entrance of Le Marais", "duration_minutes": 60, "num_players": 2}, "rules": ["Complete as many tasks as possible within 60 minutes", "Take photos or notes as proof of completion", "Stay within the designated area at all times", "No entering private buildings or restricted areas", "This game is suitable for adults"], "tasks": [{"task_id": "t1", "title": "Task 1: Explore the main square", "description": "Find and document something interesting in the main square", "location_hint": "Navigate to the main square and look for distinctive features", "points": 15, "time_limit_minutes": 8, "proof_type": "photo", "hint": "Look for signs or landmarks in the main square", "safety_note": "Stay on public paths and avoid restricted areas"}, {"task_id": "t2", "title": "Task 2: Explore the city center", "description": "Find and document something interesting in the city center", "location_hint": "Navigate to the city center and look for distinctive features", "points": 20, "time_limit_minutes": 10, "proof_type": "observation", "hint": "Look for signs or landmarks in the city center", "safety_note": "Stay on public paths and avoid restricted areas"}, {"task_id": "t3", "title": "Task 3: Explore the park area", "description": "Find and document something interesting in the park area", "location_hint": "Navigate to the park area and look for distinctive features", "points": 25, "time_limit_minutes": 12, "proof_type": "text", "hint": "Look for signs or landmarks in the park area", "safety_note": "Stay on public paths and avoid restricted areas"}, {"task_id": "t4", "title": "Task 4: Explore the landmark district", "description": "Find and document something interesting in the landmark district", "location_hint": "Navigate to the landmark district and look for distinctive features", "points": 30, "time_limit_minutes": 14, "proof_type": "photo", "hint": "Look for signs or landmarks in the landmark district", "safety_note": "Stay on public paths and avoid restricted areas"}], "global_hints": ["Explore systematically from the meeting point outward", "Ask locals for directions if needed", "Time management is key - don't spend too long on any single task"], "score_rules": ["Each task completed: full points", "Early completion: +1 bonus point per minute under limit", "Hints used: -5 points per hint", "Late arrival at meeting point: -10 points per minute"], "tie_breaker": "Winner is the player with the most points when time expires. Ties broken by earliest completion time.", "safety": {"allowed_zone": "Le Marais", "forbidden_behaviors": ["Entering buildings without permission", "Crossing busy streets recklessly", "Approaching strangers", "Leaving the designated area"], "adult_supervision": false, "stop_conditions": ["If a player feels unsafe, the game stops immediately", "If weather becomes severe, relocate to shelter", "If anyone is injured, call emergency services"]}, "story_seed": {"tone": "playful", "motifs": ["exploration", "discovery", "teamwork"], "recap_style": "episode_recap"}}, "adventure_code": "JWR5RV", "teams": ["team-a", "team-b"], "players": [{"name": "Player 1", "team_id": "team-a"}], "photos": []}, "32e1bbe3-50a1-4e43-a7c2-8ecf3f781550": {"config": {"game_type": "scavenger_hunt", "city": "Paris", "area": "Le Marais", "location_type": "mixed", "duration_minutes": 60, "num_players": 4, "difficulty": "medium", "age_group": "adults", "energy_level": "medium", "photo_enabled": true}, "game": {"game_id": "game_042", "title": "Le Marais Mystery Hunt", "theme": "Scavenger Hunt", "setup": {"city": "Paris", "area": "Le Marais", "meeting_point": "Place des Vosges", "duration_minutes": 60, "num_players": 4}, "rules": ["Stay within Le Marais boundaries", "No entering buildings or private property", "No proximity to water, traffic, or rail lines", "No interacting with strangers", "All tasks must be observed, no purchases"], "tasks": [{"task_id": "t1", "title": "Historic Fountain", "description": "Observe the historic fountain at Place des Vosges.", "location_hint": "Place des Vosges", "points": 20, "time_limit_minutes": null, "proof_type": "observation", "hint": "Look for the historic fountain", "safety_note": "Only observe, no touching or entering"}, {"task_id": "t2", "title": "Colorful Door", "description": "Observe the colorful door on Rue des Rosiers.", "location_hint": "Rue des Rosiers", "points": 15, "time_limit_minutes": null, "proof_type": "observation", "hint": "Look for the colorful door", "safety_note": "Only observe, no touching or entering"}, {"task_id": "t3", "title": "Task 3", "description": "Explore and discover something interesting in the area", "location_hint": "Look around the central area", "points": 25, "time_limit_minutes": null, "proof_type": "text", "hint": "Check visible landmarks and signs for clues", "safety_note": "Stay on public paths and respect local rules"}], "global_hints": ["Le Marais is a historic district", "All tasks are observation-based"], "score_rules": ["Higher points win", "Tie broken by total points"], "tie_breaker": "Highest total points wins", "safety": {"allowed_zone": "Le Marais public streets and parks", "forbidden_behaviors": ["Entering buildings or private property", "Proximity to water, traffic, or rail lines", "Interacting with strangers", "Purchasing items"], "adult_supervision": false, "stop_conditions": ["All tasks completed", "Time limit 60 minutes reached", "Any rule violation"]}, "story_seed": {"tone": "playful", "motifs": ["historic architecture", "urban exploration"], "recap_style": "Summarize each player's observations and points"}}, "adventure_code": "NKPZKT", "teams": ["team-a", "team-b"], "players": [{"name": "Nithin", "team_id": "team-a"}], "photos": []}, "4b03fbfc-4326-4325-8975-e4216c44e03e": {"config": {"game_type": "scavenger_hunt", "city": "Paris", "area": "Le Marais", "location_type": "mixed", "duration_minutes": 60, "num_players": 4, "difficulty": "medium", "age_group": "adults", "energy_level": "medium", "photo_enabled": true}, "game": {"game_id": "game_042", "title": "Le Marais Mystery Hunt", "theme": "Observational scavenger hunt", "setup": {"city": "Paris", "area": "Le Marais", "meeting_point": "Place des Vosges", "duration_minutes": 60, "num_players": 4}, "rules": ["Stay within Le Marais boundaries", "No entering buildings or private property", "No crossing traffic or water", "No interacting with strangers", "No purchases", "All tasks are observation only"], "tasks": [{"task_id": "t1", "title": "Vosges View", "description": "Find the historic square Place des Vosges and note its symmetrical layout.", "location_hint": "Place des Vosges", "points": 20, "time_limit_minutes": null, "proof_type": "observation", "hint": "Look for the central fountain", "safety_note": "Stay on pavement, no building entry"}, {"task_id": "t2", "title": "Rue des Rosiers", "description": "Observe the vibrant street market stalls and colorful awnings.", "location_hint": "Rue des Rosiers", "points": 15, "time_limit_minutes": null, "proof_type": "observation", "hint": "Notice the flower displays", "safety_note": "No entering stalls"}, {"task_id": "t3", "title": "H\u00f4tel de Ville View", "description": "Observe the fa\u00e7ade of the H\u00f4tel de Ville from the street.", "location_hint": "Place des Vosges near H\u00f4tel de Ville", "points": 20, "time_limit_minutes": null, "proof_type": "observation", "hint": "Look for the clock tower", "safety_note": "No entering building"}], "global_hints": ["Stay within Le Marais", "Use observation only", "No purchases"], "score_rules": ["Highest points wins", "All tasks completed for bonus"], "tie_breaker": "First player to reach 50 points wins", "safety": {"allowed_zone": "Le Marais streets and public squares", "forbidden_behaviors": ["Entering buildings", "Crossing traffic", "Approaching water", "Interacting with strangers", "Purchasing items"], "adult_supervision": true, "stop_conditions": ["Time limit 60 minutes", "Safety breach", "All tasks completed"]}, "story_seed": {"tone": "playful", "motifs": ["hidden clues", "Parisian charm", "observation"], "recap_style": "Gather at the meeting point for a recap of findings."}}, "adventure_code": "NXWHQK", "teams": ["team-a", "team-b"], "players": [{"name": "Nithin", "team_id": "team-a"}], "photos": []}}}
app/schemas/journal_schema.json CHANGED
@@ -52,6 +52,38 @@
52
  "type": "string"
53
  },
54
  "description": "References to associated photos"
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
55
  }
56
  }
57
  }
 
52
  "type": "string"
53
  },
54
  "description": "References to associated photos"
55
+ },
56
+ "transcript_source": {
57
+ "type": "string",
58
+ "enum": ["typed", "asr", "hybrid"],
59
+ "description": "Where the transcript text came from: typed by the user, produced by the ASR service, or hybrid (ASR + manual edits)."
60
+ },
61
+ "audio_ref": {
62
+ "type": "string",
63
+ "description": "Path or URL of the recorded audio clip associated with this entry."
64
+ },
65
+ "asr": {
66
+ "type": "object",
67
+ "description": "Metadata describing the ASR pass that produced the transcript (when transcript_source is 'asr' or 'hybrid').",
68
+ "properties": {
69
+ "model": {
70
+ "type": "string",
71
+ "description": "ASR model identifier (e.g. CohereLabs/cohere-transcribe-03-2026)."
72
+ },
73
+ "language": {
74
+ "type": "string",
75
+ "description": "Language code used for transcription."
76
+ },
77
+ "status": {
78
+ "type": "string",
79
+ "enum": ["ok", "skipped", "error"],
80
+ "description": "Outcome of the ASR pass."
81
+ },
82
+ "error": {
83
+ "type": ["string", "null"],
84
+ "description": "Error message if status is 'error' or 'skipped'."
85
+ }
86
+ }
87
  }
88
  }
89
  }
app/services/asr.py ADDED
@@ -0,0 +1,269 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Automatic Speech Recognition (ASR) module.
2
+
3
+ Default model: ``CohereLabs/cohere-transcribe-03-2026`` (Apache 2.0, 2B
4
+ parameters, 14 languages).
5
+
6
+ Uses the official πŸ€— Transformers integration:
7
+
8
+ - transformers >= 5.4.0
9
+ - torch
10
+ - huggingface_hub
11
+ - soundfile, librosa
12
+ - sentencepiece, protobuf
13
+
14
+ Audio is auto-resampled to 16 kHz and downmixed to mono by the
15
+ processor. The model supports 14 languages (English, French, German,
16
+ Italian, Spanish, Portuguese, Greek, Dutch, Polish, Mandarin, Japanese,
17
+ Korean, Vietnamese, Arabic) but does not auto-detect language, so
18
+ callers should pass an explicit language code.
19
+
20
+ Like the other model services, the model is **lazy-loaded** so it does
21
+ not block app startup. Setting the environment variable
22
+ ``CITYQUEST_SKIP_MODEL=1`` (or ``CITYQUEST_FAST_TEST=1``) will make all
23
+ ASR calls return an empty string with a warning so the rest of the
24
+ pipeline (typed journal fallback, scoring, recap) keeps working.
25
+ """
26
+
27
+ from __future__ import annotations
28
+
29
+ import os
30
+ from pathlib import Path
31
+ from typing import Optional
32
+
33
+
34
+ # ── Configuration ───────────────────────────────────────────────────────────
35
+ MODEL_ID = "CohereLabs/cohere-transcribe-03-2026"
36
+ SUPPORTED_LANGUAGES = [
37
+ "en", "fr", "de", "it", "es", "pt", "el", "nl", "pl",
38
+ "zh", "ja", "ko", "vi", "ar",
39
+ ]
40
+ DEFAULT_LANGUAGE = "en"
41
+ MAX_NEW_TOKENS = 256
42
+ SAMPLING_RATE = 16_000
43
+
44
+ # Skipped under CITYQUEST_SKIP_MODEL (model download) and CITYQUEST_FAST_TEST
45
+ # (used by `test_end_to_end.py`).
46
+ _skip_env_var = "CITYQUEST_SKIP_MODEL"
47
+ _fast_test_env_var = "CITYQUEST_FAST_TEST"
48
+
49
+
50
+ # ── Lazy model state ────────────────────────────────────────────────────────
51
+ _processor = None
52
+ _model = None
53
+ _model_loaded = False
54
+ _load_attempted = False
55
+
56
+
57
+ def _is_skipped() -> bool:
58
+ """True if ASR model loading/inference should be short-circuited."""
59
+ return bool(os.environ.get(_skip_env_var) or os.environ.get(_fast_test_env_var))
60
+
61
+
62
+ def _load_model() -> bool:
63
+ """Lazy-load the Cohere Transcribe model + processor on first call.
64
+
65
+ Returns:
66
+ ``True`` if the model is ready, ``False`` if it was skipped or
67
+ failed to load.
68
+ """
69
+ global _processor, _model, _model_loaded, _load_attempted
70
+ if _model_loaded:
71
+ return _processor is not None and _model is not None
72
+ _load_attempted = True
73
+
74
+ if _is_skipped():
75
+ print("[asr] skip env var set β€” ASR model not loaded")
76
+ return False
77
+
78
+ try:
79
+ import torch
80
+ from transformers import AutoProcessor, CohereAsrForConditionalGeneration
81
+ from huggingface_hub import login
82
+
83
+ # Try logging in if a token is provided in env (gated model).
84
+ token = os.environ.get("HF_TOKEN") or os.environ.get("HUGGINGFACE_HUB_TOKEN")
85
+ if token:
86
+ try:
87
+ login(token=token, add_to_git_credential=False)
88
+ except Exception as exc:
89
+ print(f"[asr] HF login skipped: {exc}")
90
+
91
+ device = "cuda" if torch.cuda.is_available() else "cpu"
92
+ dtype = torch.bfloat16 if device == "cuda" else torch.float32
93
+
94
+ print(f"[asr] Loading {MODEL_ID} on {device} ({dtype}) …")
95
+ _processor = AutoProcessor.from_pretrained(MODEL_ID)
96
+ _model = CohereAsrForConditionalGeneration.from_pretrained(
97
+ MODEL_ID, torch_dtype=dtype, device_map="auto",
98
+ )
99
+ _model.eval()
100
+ _model_loaded = True
101
+ print("[asr] Model ready")
102
+ return True
103
+
104
+ except ImportError as exc:
105
+ print(
106
+ f"[asr] Required packages not installed: {exc}. "
107
+ "Install with: pip install 'transformers>=5.4.0' torch "
108
+ "huggingface_hub soundfile librosa sentencepiece protobuf"
109
+ )
110
+ except Exception as exc:
111
+ print(f"[asr] Model load failed: {type(exc).__name__}: {exc}")
112
+ return False
113
+
114
+
115
+ def _load_audio(audio_path: str) -> Optional[list[float]]:
116
+ """Load an audio file as a mono float32 waveform at 16 kHz.
117
+
118
+ Tries multiple loading strategies in order:
119
+ 1. ``transformers.audio_utils.load_audio`` (preferred, handles
120
+ resampling + mono conversion).
121
+ 2. ``soundfile`` fallback.
122
+ """
123
+ if not Path(audio_path).exists():
124
+ print(f"[asr] Audio file not found: {audio_path}")
125
+ return None
126
+
127
+ try:
128
+ from transformers.audio_utils import load_audio
129
+
130
+ waveform = load_audio(audio_path, sampling_rate=SAMPLING_RATE)
131
+ return list(waveform)
132
+ except ImportError:
133
+ pass
134
+ except Exception as exc:
135
+ print(f"[asr] transformers audio load failed: {exc}")
136
+
137
+ # Fallback: soundfile + numpy
138
+ try:
139
+ import numpy as np
140
+ import soundfile as sf
141
+
142
+ data, sr = sf.read(audio_path, dtype="float32", always_2d=False)
143
+ if data.ndim > 1:
144
+ data = data.mean(axis=1)
145
+ if sr != SAMPLING_RATE:
146
+ # simple linear resample β€” good enough for short voice notes
147
+ ratio = SAMPLING_RATE / sr
148
+ new_len = int(len(data) * ratio)
149
+ data = np.interp(
150
+ np.linspace(0, len(data), new_len, endpoint=False),
151
+ np.arange(len(data)),
152
+ data,
153
+ ).astype("float32")
154
+ return list(data)
155
+ except ImportError:
156
+ print(
157
+ "[asr] soundfile not installed β€” install with: "
158
+ "pip install soundfile librosa"
159
+ )
160
+ except Exception as exc:
161
+ print(f"[asr] soundfile load failed: {exc}")
162
+ return None
163
+
164
+
165
+ # ── Public API ──────────────────────────────────────────────────────────────
166
+
167
+ def transcribe(
168
+ audio_path: str,
169
+ language: str = DEFAULT_LANGUAGE,
170
+ max_new_tokens: int = MAX_NEW_TOKENS,
171
+ ) -> dict:
172
+ """Transcribe a voice journal audio clip to text.
173
+
174
+ Args:
175
+ audio_path: Path to an audio file (wav/mp3/m4a/ogg/webm).
176
+ language: One of ``SUPPORTED_LANGUAGES``. Defaults to ``"en"``.
177
+ max_new_tokens: Maximum number of tokens to generate.
178
+
179
+ Returns:
180
+ Dict with keys ``transcript`` (str), ``language`` (str),
181
+ ``model`` (str), ``status`` (str β€” ``"ok"`` | ``"skipped"`` |
182
+ ``"error"``), and ``error`` (str | None). On non-ok statuses
183
+ ``transcript`` is ``""``.
184
+ """
185
+ if language not in SUPPORTED_LANGUAGES:
186
+ print(
187
+ f"[asr] Language '{language}' not in supported set β€” "
188
+ f"falling back to '{DEFAULT_LANGUAGE}'."
189
+ )
190
+ language = DEFAULT_LANGUAGE
191
+
192
+ if _is_skipped():
193
+ return {
194
+ "transcript": "",
195
+ "language": language,
196
+ "model": MODEL_ID,
197
+ "status": "skipped",
198
+ "error": "ASR model skipped (CITYQUEST_SKIP_MODEL/CITYQUEST_FAST_TEST).",
199
+ }
200
+
201
+ if not _load_model():
202
+ return {
203
+ "transcript": "",
204
+ "language": language,
205
+ "model": MODEL_ID,
206
+ "status": "error",
207
+ "error": "ASR model unavailable.",
208
+ }
209
+
210
+ waveform = _load_audio(audio_path)
211
+ if waveform is None:
212
+ return {
213
+ "transcript": "",
214
+ "language": language,
215
+ "model": MODEL_ID,
216
+ "status": "error",
217
+ "error": f"Could not load audio at {audio_path}.",
218
+ }
219
+
220
+ try:
221
+ import torch
222
+
223
+ inputs = _processor(
224
+ waveform,
225
+ sampling_rate=SAMPLING_RATE,
226
+ return_tensors="pt",
227
+ language=language,
228
+ )
229
+ inputs = {
230
+ k: v.to(_model.device, dtype=_model.dtype) if hasattr(v, "to") else v
231
+ for k, v in inputs.items()
232
+ }
233
+
234
+ with torch.inference_mode():
235
+ outputs = _model.generate(
236
+ **inputs,
237
+ max_new_tokens=max_new_tokens,
238
+ )
239
+ text = _processor.decode(outputs[0], skip_special_tokens=True).strip()
240
+
241
+ return {
242
+ "transcript": text,
243
+ "language": language,
244
+ "model": MODEL_ID,
245
+ "status": "ok",
246
+ "error": None,
247
+ }
248
+
249
+ except Exception as exc:
250
+ print(f"[asr] Inference failed: {type(exc).__name__}: {exc}")
251
+ return {
252
+ "transcript": "",
253
+ "language": language,
254
+ "model": MODEL_ID,
255
+ "status": "error",
256
+ "error": f"{type(exc).__name__}: {exc}",
257
+ }
258
+
259
+
260
+ def transcribe_text(
261
+ audio_path: str,
262
+ language: str = DEFAULT_LANGUAGE,
263
+ ) -> str:
264
+ """Convenience wrapper that returns just the transcript string.
265
+
266
+ Returns ``""`` on any failure so callers can fall back to typed
267
+ journal input without exception handling.
268
+ """
269
+ return transcribe(audio_path, language=language).get("transcript", "")
app/services/journal.py CHANGED
@@ -62,16 +62,25 @@ LOCATION_KEYWORDS = [
62
 
63
  # ── Transcription ───────────────────────────────────────────────────────────
64
 
65
- def transcribe_journal(audio_path: str) -> str:
 
 
 
 
66
  """Transcribe voice journal audio to text.
67
 
68
- Tries (in order):
69
- 1. ``whisper`` Python package (OpenAI Whisper, runs locally).
70
- 2. Returns an empty string with a warning so downstream code
71
- still works and the user can fall back to typed input.
 
 
 
72
 
73
  Args:
74
  audio_path: Path to a recorded audio file (wav/mp3/m4a/ogg/webm).
 
 
75
 
76
  Returns:
77
  Transcribed text, or ``""`` if transcription is unavailable.
@@ -80,7 +89,24 @@ def transcribe_journal(audio_path: str) -> str:
80
  if not path.exists():
81
  raise FileNotFoundError(f"Audio file not found: {audio_path}")
82
 
83
- # Try local Whisper
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
84
  try:
85
  import whisper # type: ignore
86
 
@@ -224,6 +250,9 @@ def create_journal_entry(
224
  task_id: Optional[str] = None,
225
  location_note: str = "",
226
  photo_refs: Optional[list[str]] = None,
 
 
 
227
  ) -> dict:
228
  """Build a complete journal entry dict matching the journal schema.
229
 
@@ -234,6 +263,12 @@ def create_journal_entry(
234
  task_id: Optional associated task ID.
235
  location_note: Where the entry was recorded.
236
  photo_refs: List of photo identifiers to attach.
 
 
 
 
 
 
237
 
238
  Returns:
239
  Full journal entry dict.
@@ -249,11 +284,18 @@ def create_journal_entry(
249
  "mood": mood,
250
  "location_note": location_note or "Unknown location",
251
  "photo_refs": photo_refs or [],
 
252
  }
253
 
254
  if task_id:
255
  entry["task_id"] = task_id
256
 
 
 
 
 
 
 
257
  return entry
258
 
259
 
 
62
 
63
  # ── Transcription ───────────────────────────────────────────────────────────
64
 
65
+ def transcribe_journal(
66
+ audio_path: str,
67
+ language: str = "en",
68
+ prefer: str = "cohere",
69
+ ) -> str:
70
  """Transcribe voice journal audio to text.
71
 
72
+ Tries (in order, controlled by ``prefer``):
73
+ 1. ``"cohere"`` (default) β€” ``CohereLabs/cohere-transcribe-03-2026``
74
+ via ``app.services.asr``. Lazy-loaded; honors the
75
+ ``CITYQUEST_SKIP_MODEL`` / ``CITYQUEST_FAST_TEST`` env vars.
76
+ 2. ``"whisper"`` β€” OpenAI Whisper, runs locally.
77
+ 3. Returns an empty string with a warning so the caller can fall
78
+ back to typed input.
79
 
80
  Args:
81
  audio_path: Path to a recorded audio file (wav/mp3/m4a/ogg/webm).
82
+ language: Language code passed to the ASR model (e.g. ``"en"``).
83
+ prefer: ``"cohere"`` | ``"whisper"`` | ``"any"``.
84
 
85
  Returns:
86
  Transcribed text, or ``""`` if transcription is unavailable.
 
89
  if not path.exists():
90
  raise FileNotFoundError(f"Audio file not found: {audio_path}")
91
 
92
+ # ── 1. Cohere Transcribe (preferred) ──────────────────────────────────
93
+ if prefer in ("cohere", "any"):
94
+ try:
95
+ from app.services.asr import transcribe_text as _cohere_transcribe
96
+
97
+ text = _cohere_transcribe(str(path), language=language)
98
+ if text:
99
+ return text
100
+ except ImportError as exc:
101
+ print(f"[journal] Cohere ASR import failed: {exc}")
102
+ except Exception as exc:
103
+ print(f"[journal] Cohere ASR failed: {type(exc).__name__}: {exc}")
104
+
105
+ if prefer == "cohere":
106
+ # Don't fall back to Whisper unless the caller asks for it.
107
+ return ""
108
+
109
+ # ── 2. Whisper fallback (opt-in) ──────────────────────────────────────
110
  try:
111
  import whisper # type: ignore
112
 
 
250
  task_id: Optional[str] = None,
251
  location_note: str = "",
252
  photo_refs: Optional[list[str]] = None,
253
+ audio_ref: Optional[str] = None,
254
+ asr_metadata: Optional[dict] = None,
255
+ transcript_source: str = "typed",
256
  ) -> dict:
257
  """Build a complete journal entry dict matching the journal schema.
258
 
 
263
  task_id: Optional associated task ID.
264
  location_note: Where the entry was recorded.
265
  photo_refs: List of photo identifiers to attach.
266
+ audio_ref: Optional path/URL of the recorded audio clip.
267
+ asr_metadata: Optional dict describing the ASR pass that
268
+ produced ``transcript`` (model, language, status, error).
269
+ transcript_source: One of ``"typed"`` | ``"asr"`` | ``"hybrid"`` β€”
270
+ whether the transcript came from typing, the ASR service,
271
+ or a combination (ASR with manual edits).
272
 
273
  Returns:
274
  Full journal entry dict.
 
284
  "mood": mood,
285
  "location_note": location_note or "Unknown location",
286
  "photo_refs": photo_refs or [],
287
+ "transcript_source": transcript_source,
288
  }
289
 
290
  if task_id:
291
  entry["task_id"] = task_id
292
 
293
+ if audio_ref:
294
+ entry["audio_ref"] = audio_ref
295
+
296
+ if asr_metadata:
297
+ entry["asr"] = asr_metadata
298
+
299
  return entry
300
 
301
 
requirements.txt CHANGED
@@ -11,4 +11,9 @@ jsonschema
11
  diffusers[torch] # uncomment for offline FLUX poster generation
12
  transformers
13
  accelerate
 
 
 
 
 
14
  # modal # uncomment to deploy to Modal cloud GPU
 
11
  diffusers[torch] # uncomment for offline FLUX poster generation
12
  transformers
13
  accelerate
14
+ # Cohere Transcribe ASR (voice journal) β€” gated model, requires HF_TOKEN
15
+ soundfile
16
+ librosa
17
+ sentencepiece
18
+ protobuf
19
  # modal # uncomment to deploy to Modal cloud GPU
test_asr.py ADDED
@@ -0,0 +1,249 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Tests for the ASR-backed journal pipeline.
2
+
3
+ These tests cover the fast/skip-model path and the metadata path so the
4
+ rest of the journal pipeline (scoring, recap, persistence) keeps
5
+ working even when no audio model is available.
6
+
7
+ Run with:
8
+
9
+ CITYQUEST_FAST_TEST=1 python test_asr.py
10
+
11
+ To exercise the real model, run without the skip env var on a GPU box.
12
+ """
13
+
14
+ from __future__ import annotations
15
+
16
+ import json
17
+ import os
18
+ import sys
19
+ import tempfile
20
+ import uuid
21
+ from pathlib import Path
22
+
23
+
24
+ # ── Fast-test guard ───────────────────────────────────────────────────────
25
+ if os.environ.get("CITYQUEST_FAST_TEST") or not os.environ.get("CITYQUEST_RUN_ASR"):
26
+ os.environ.setdefault("CITYQUEST_SKIP_MODEL", "1")
27
+
28
+ passed = 0
29
+ failed = 0
30
+ errors: list[str] = []
31
+
32
+
33
+ def check(label: str, condition: bool, detail: str = "") -> None:
34
+ global passed, failed
35
+ if condition:
36
+ passed += 1
37
+ print(f" βœ“ PASS: {label}")
38
+ else:
39
+ failed += 1
40
+ msg = f"βœ— FAIL: {label} β€” {detail}" if detail else f"βœ— FAIL: {label}"
41
+ errors.append(msg)
42
+ print(f" {msg}")
43
+
44
+
45
+ def main() -> int:
46
+ global passed, failed
47
+ print("=" * 80)
48
+ print("ASR JOURNAL PIPELINE β€” TESTS")
49
+ print("=" * 80)
50
+
51
+ # ── T1: ASR module imports & constants ──────────────────────────────
52
+ print("\n" + "-" * 80)
53
+ print("T1: ASR MODULE IMPORTS")
54
+ print("-" * 80)
55
+ try:
56
+ from app.services.asr import (
57
+ MODEL_ID,
58
+ SUPPORTED_LANGUAGES,
59
+ DEFAULT_LANGUAGE,
60
+ transcribe,
61
+ transcribe_text,
62
+ )
63
+ check("ASR module imports", True)
64
+ check("Model id is Cohere Transcribe", "cohere-transcribe" in MODEL_ID,
65
+ f"Got {MODEL_ID}")
66
+ check("Supports 14 languages", len(SUPPORTED_LANGUAGES) == 14,
67
+ f"Got {len(SUPPORTED_LANGUAGES)}")
68
+ check("Default language is English", DEFAULT_LANGUAGE == "en")
69
+ except Exception as e:
70
+ check("ASR module imports", False, str(e))
71
+ _summary()
72
+ return 1
73
+
74
+ # ── T2: transcribe() returns structured result ──────────────────────
75
+ print("\n" + "-" * 80)
76
+ print("T2: transcribe() CONTRACT")
77
+ print("-" * 80)
78
+ try:
79
+ # Use a fake file path β€” skipped mode should never reach the disk
80
+ result = transcribe("/nonexistent/audio.wav", language="en")
81
+ check("transcribe() returns dict", isinstance(result, dict))
82
+ for key in ("transcript", "language", "model", "status", "error"):
83
+ check(f"transcribe() has '{key}'", key in result)
84
+ check("status is 'ok' | 'skipped' | 'error'",
85
+ result["status"] in ("ok", "skipped", "error"),
86
+ f"Got {result['status']}")
87
+ check("transcript is a string", isinstance(result["transcript"], str))
88
+ check("model is reported", bool(result["model"]))
89
+ except Exception as e:
90
+ check("transcribe() contract", False, str(e))
91
+
92
+ # ── T3: Skipped mode is silent and empty ────────────────────────────
93
+ print("\n" + "-" * 80)
94
+ print("T3: SKIP-MODE BEHAVIOR")
95
+ print("-" * 80)
96
+ try:
97
+ os.environ["CITYQUEST_SKIP_MODEL"] = "1"
98
+ result = transcribe("/tmp/fake.wav", language="en")
99
+ check("Skip mode returns empty transcript", result["transcript"] == "")
100
+ check("Skip mode reports status=skipped", result["status"] == "skipped")
101
+ # Restore to be safe
102
+ del os.environ["CITYQUEST_SKIP_MODEL"]
103
+ except Exception as e:
104
+ check("Skip-mode behavior", False, str(e))
105
+
106
+ # ── T4: transcribe_journal() returns empty on skip ──────────────────
107
+ print("\n" + "-" * 80)
108
+ print("T4: JOURNAL.transcribe_journal() FALLBACK")
109
+ print("-" * 80)
110
+ try:
111
+ os.environ["CITYQUEST_SKIP_MODEL"] = "1"
112
+ from app.services.journal import transcribe_journal
113
+
114
+ # Create a tiny valid wav so path checks pass
115
+ with tempfile.NamedTemporaryFile(suffix=".wav", delete=False) as fh:
116
+ wav_path = fh.name
117
+ # Minimal RIFF header β€” not a real wav, but enough for the
118
+ # existence check in transcribe_journal()
119
+ fh.write(b"RIFF\x24\x00\x00\x00WAVEfmt ")
120
+
121
+ try:
122
+ text = transcribe_journal(wav_path, language="en")
123
+ check("transcribe_journal returns str", isinstance(text, str))
124
+ check("transcribe_journal empty when ASR skipped", text == "",
125
+ f"Got {text!r}")
126
+ finally:
127
+ try:
128
+ os.unlink(wav_path)
129
+ except OSError:
130
+ pass
131
+ except Exception as e:
132
+ check("Journal.transcribe_journal fallback", False, str(e))
133
+
134
+ # ── T5: create_journal_entry() supports ASR metadata ───────────────
135
+ print("\n" + "-" * 80)
136
+ print("T5: JOURNAL ENTRY METADATA")
137
+ print("-" * 80)
138
+ try:
139
+ from app.services.journal import create_journal_entry, save_journal_entry, load_journal_entries
140
+
141
+ asr_meta = {
142
+ "model": "CohereLabs/cohere-transcribe-03-2026",
143
+ "language": "en",
144
+ "status": "ok",
145
+ "error": None,
146
+ }
147
+ entry = create_journal_entry(
148
+ transcript="We just found the mural near the canal, amazing!",
149
+ session_id="test-asr",
150
+ team_id="team-a",
151
+ task_id="t1",
152
+ location_note="Canal area",
153
+ audio_ref="/tmp/clip.wav",
154
+ asr_metadata=asr_meta,
155
+ transcript_source="asr",
156
+ )
157
+ check("Entry has transcript_source", entry.get("transcript_source") == "asr")
158
+ check("Entry has audio_ref", entry.get("audio_ref") == "/tmp/clip.wav")
159
+ check("Entry has asr metadata", isinstance(entry.get("asr"), dict))
160
+ check("ASR metadata has model", entry["asr"]["model"] == asr_meta["model"])
161
+ check("ASR metadata has language", entry["asr"]["language"] == "en")
162
+ check("ASR metadata has status", entry["asr"]["status"] == "ok")
163
+
164
+ # Schema validation (optional β€” only if jsonschema is installed)
165
+ try:
166
+ import jsonschema # type: ignore
167
+ schema_path = Path("app/schemas/journal_schema.json")
168
+ schema = json.loads(schema_path.read_text())
169
+ jsonschema.validate(instance=entry, schema=schema)
170
+ check("Entry validates against journal_schema.json", True)
171
+ except ImportError:
172
+ check("Entry validates against journal_schema.json (skipped, no jsonschema)", True)
173
+
174
+ # Persistence round-trip
175
+ save_journal_entry(entry)
176
+ loaded = load_journal_entries(session_id="test-asr")
177
+ check("Entry persisted with ASR metadata", len(loaded) >= 1)
178
+ if loaded:
179
+ check("Loaded entry has transcript_source",
180
+ loaded[-1].get("transcript_source") == "asr")
181
+ check("Loaded entry has asr block",
182
+ isinstance(loaded[-1].get("asr"), dict))
183
+ except Exception as e:
184
+ check("Journal entry metadata", False, str(e))
185
+
186
+ # ── T6: app.record_journal() handles voice path ─────────────────────
187
+ print("\n" + "-" * 80)
188
+ print("T6: app.record_journal() VOICE PATH (skip-mode)")
189
+ print("-" * 80)
190
+ try:
191
+ os.environ["CITYQUEST_SKIP_MODEL"] = "1"
192
+ # We can't import the whole app (it boots Gradio), so test the
193
+ # logic by directly exercising the journal functions.
194
+ from app.services.journal import transcribe_journal, create_journal_entry
195
+ from app.services.asr import transcribe as _asr_transcribe
196
+
197
+ with tempfile.NamedTemporaryFile(suffix=".wav", delete=False) as fh:
198
+ wav_path = fh.name
199
+ fh.write(b"RIFF\x24\x00\x00\x00WAVEfmt ")
200
+
201
+ try:
202
+ asr_result = _asr_transcribe(wav_path, language="en")
203
+ check("ASR returns skipped status", asr_result["status"] == "skipped")
204
+ check("ASR transcript is empty", asr_result["transcript"] == "")
205
+
206
+ # When ASR produces no text, journal creation should still
207
+ # be possible via the typed/manual correction path.
208
+ entry = create_journal_entry(
209
+ transcript="Typed correction (ASR was unavailable).",
210
+ session_id="test-asr-voice",
211
+ audio_ref=wav_path,
212
+ asr_metadata={
213
+ "model": asr_result["model"],
214
+ "language": asr_result["language"],
215
+ "status": asr_result["status"],
216
+ "error": asr_result["error"],
217
+ },
218
+ transcript_source="typed",
219
+ )
220
+ check("Hybrid entry created with asr metadata",
221
+ entry.get("asr", {}).get("status") == "skipped")
222
+ finally:
223
+ try:
224
+ os.unlink(wav_path)
225
+ except OSError:
226
+ pass
227
+ except Exception as e:
228
+ check("Voice path in skip-mode", False, str(e))
229
+
230
+ _summary()
231
+ return 0 if failed == 0 else 1
232
+
233
+
234
+ def _summary() -> None:
235
+ total = passed + failed
236
+ print("\n" + "=" * 80)
237
+ if failed == 0:
238
+ print(f"RESULTS: {passed}/{total} tests passed β€” ALL CLEAR πŸŽ‰")
239
+ else:
240
+ print(f"RESULTS: {passed}/{total} tests passed β€” {failed} FAILED")
241
+ print("=" * 80)
242
+ if errors:
243
+ print("\nFailed tests:")
244
+ for e in errors:
245
+ print(f" {e}")
246
+
247
+
248
+ if __name__ == "__main__":
249
+ sys.exit(main())