Dataset Viewer
Auto-converted to Parquet Duplicate
Search is not available for this dataset
participant_id
int64
361
467
S2_P1
int64
1
5
F2_P1
int64
1
5
C1_P1
int64
1
5
A1_P1
int64
1
5
M1_P1
int64
1
5
V1_P1
int64
1
5
S1_P1
int64
1
5
G2_P1
int64
1
5
A2_P1
int64
1
5
B2_P1
int64
1
5
361
2
2
2
4
2
2
2
2
4
1
362
1
1
2
2
2
3
2
3
1
1
363
1
1
2
1
1
1
4
1
2
5
364
1
1
1
4
1
1
5
2
4
1
365
2
2
2
1
1
1
5
4
4
4
366
4
1
5
1
4
1
1
1
1
1
367
1
2
5
1
1
4
1
4
1
4
368
1
4
1
1
1
1
1
1
4
1
369
1
2
2
1
4
1
1
2
1
1
370
1
2
5
1
1
1
1
1
1
1
371
1
1
3
1
4
2
1
1
1
1
372
1
1
1
2
2
1
2
1
1
1
373
1
2
2
1
5
1
1
2
1
1
374
1
2
5
3
3
3
3
3
5
3
375
1
2
2
2
4
2
4
1
1
1
376
2
2
2
1
2
1
1
1
1
2
377
2
1
1
4
1
1
1
1
1
4
378
4
2
1
1
5
1
4
1
5
4
379
3
2
1
2
1
1
1
5
2
4
380
1
1
2
3
3
3
5
4
5
1
381
2
4
5
3
1
4
1
2
1
4
382
4
1
1
5
5
4
4
4
5
2
383
4
4
5
4
4
4
1
4
4
4
384
1
2
2
4
4
1
1
4
4
4
385
4
1
1
5
5
5
1
1
5
4
386
1
1
1
1
4
1
4
4
5
3
387
5
5
5
5
4
4
5
4
4
4
388
4
1
4
4
1
4
4
5
5
5
389
4
2
1
2
1
1
4
1
4
4
390
4
5
5
4
4
3
5
4
4
5
391
4
1
5
1
4
4
4
5
5
4
392
5
5
5
4
4
4
4
2
1
4
393
2
4
4
4
4
2
4
4
4
4
394
2
4
5
2
2
2
4
4
4
4
395
4
4
4
2
4
2
2
2
4
4
396
5
5
1
1
4
2
4
5
5
3
397
1
1
1
1
1
1
1
5
5
3
398
4
2
4
4
4
4
4
4
4
3
399
2
3
4
1
5
2
4
5
5
4
400
2
4
2
2
4
2
2
4
4
4
401
2
4
4
5
4
2
2
4
4
3
402
2
4
2
2
4
2
2
4
4
4
403
4
3
1
5
5
5
5
5
3
4
404
2
4
4
2
4
2
2
2
4
4
405
4
3
2
4
5
5
5
4
1
4
406
2
4
2
2
2
2
2
4
4
3
407
5
2
5
5
4
4
4
2
1
3
408
1
5
5
5
5
1
2
5
5
4
409
4
4
2
2
4
4
5
4
1
2
410
2
2
2
2
4
2
1
4
2
4
413
1
3
2
5
5
2
5
5
2
4
414
2
5
2
2
2
2
2
5
5
4
415
2
4
4
2
4
2
2
4
4
4
416
2
3
5
2
4
2
2
4
4
4
417
2
2
4
2
2
2
2
4
4
4
419
2
4
2
2
4
4
2
4
4
4
420
2
4
4
4
4
3
2
4
4
4
421
2
3
4
4
2
4
2
4
4
3
422
2
4
2
4
4
2
2
4
4
4
423
2
4
4
4
2
2
2
2
4
3
424
4
4
2
2
4
2
2
4
4
3
425
1
4
1
2
4
3
1
4
4
4
426
2
5
1
1
5
1
1
5
5
5
427
2
5
2
1
5
1
1
5
5
5
428
4
4
5
2
4
4
2
4
4
4
429
2
5
2
2
2
1
1
1
5
2
432
2
4
4
2
4
4
2
2
4
3
446
2
3
4
2
2
2
2
4
4
3
447
2
4
1
1
4
2
1
5
5
5
448
2
2
5
1
2
1
1
1
5
5
449
1
2
1
1
5
3
2
5
5
5
450
1
4
1
1
4
1
1
4
4
4
451
2
4
4
2
4
1
2
4
4
4
452
2
4
1
2
5
1
1
5
1
5
453
1
4
1
1
4
2
1
5
4
4
454
2
2
5
1
4
1
4
5
4
4
455
2
1
1
1
5
1
2
5
4
5
457
2
3
4
2
4
2
2
4
2
3
458
2
4
1
2
5
3
2
5
4
4
459
1
1
1
5
5
1
5
1
1
4
460
1
1
1
4
4
4
4
4
4
4
461
3
5
1
1
4
1
4
1
4
1
463
5
1
1
5
5
4
4
1
4
1
467
2
2
4
2
4
3
1
3
4
3

Data release

Data from a three-condition randomized field study on women's health misinformation, with 434 low-literacy women in Mangolpuri, North West Delhi, across three sessions over one month.

Paper: arXiv:2609.19364

Contents

misinformation/
  womens_health_misinformation.csv    43 beliefs elicited from health providers
belief_ratings/
  belief_ratings_cultural.csv         188 participants x 20 items x 4 timepoints
  belief_ratings_non_cultural.csv     162 participants x 20 items x 4 timepoints
  belief_ratings_control.csv           84 participants x 10 items, 1 timepoint
  item_key.csv                         the 20 survey items
  response_scale.csv                   what the 1-5 response values mean
qualitative/
  qual_responses.csv                  350 post-video interviews
  demographics.csv                    434 participant profiles
analysis/
  run_all.py                          reproduces the three main results
  01_belief_change.py, 02_learning.py, 03_retention.py
video_generation/
  generate_clips.py, ...              the pipeline that produced the two videos
  chunks.csv                          the 93 generated clips: text, length, duration
  prompts_by_segment.csv              the prompt sent for each segment, both conditions

Code

analysis/ reproduces the study's three results directly from the belief-rating files here — immediate belief change, learning, and retention. It needs numpy, scipy, pandas and statsmodels, and nothing else:

cd analysis && python3 run_all.py

It regenerates the paper's figures exactly — the within-session gains, the baseline-adjusted advantage at each wave, the learning contrasts against control, and retention at three weeks.

Conditions

condition participant_id n saw
culturally adaptive 1–190 188 AI video, presenter matched to the viewer's community
culturally neutral 191–360 162 the same video, presenter not matched
no-video control 361–467 84 no video

How the files join

belief_ratings/item_key.csv
    item_id        ─────────►  belief_ratings_*.csv column prefix
                                   e.g. item S2  ->  S2_P1_pre, S2_P1_post, S2_P2, S2_P3

belief_ratings/belief_ratings_*.csv
    cell value 1-5 ─────────►  response_scale.value
    participant_id ─────────►  qualitative/qual_responses.participant_id
                               qualitative/demographics.participant_id

misinformation/womens_health_misinformation.csv

43 rows, 5 columns. One row per false belief reported by health providers (doctors, pharmacists, ASHA and anganwadi workers) in the study communities.

column description
misinformation the false belief, in English
verbatim_quote_deidentified the provider's own words describing it
reported_clinical_impact the harm, as the provider described it
provider_count how many providers independently reported it
in_final_data yes for the 10 beliefs carried into the study; no for the other 33

10 of the 43 beliefs were carried into the study (in_final_data = yes); the other 33 were not. The quote and clinical-impact columns are filled for 12 beliefs, which include all 10 in the final data; the remaining 31 rows record the belief and how many providers reported it.

Providers are referred to by code (P1–P11) inside the quote text. The same code means the same person throughout.


belief_ratings/

belief_ratings_cultural.csv, belief_ratings_non_cultural.csv

188 and 162 rows, 81 columns.

column description
participant_id 1–190 cultural, 191–360 non-cultural
{item}_{timepoint} the 1–5 response, 80 columns

{item} is one of the 20 ids in item_key.csv. {timepoint} is one of:

timepoint when
P1_pre session 1, before the video
P1_post session 1, immediately after the video
P2 two weeks later
P3 three weeks later

A blank cell means the item was not asked at that timepoint. Which items were asked when is set by the item's block, and is recorded in item_key.csv:

block items asked at
A S2, F2, C1, A1, M1 P1_pre, P1_post, P2, P3
B V1, S1, G2, A2, B2 P1_post, P2, P3
C D1, K1, C2, F1, M2 P2, P3
D V2, K2, D2, B1, G1 P3

So 5 items are answered at P1_pre, 10 at P1_post, 15 at P2, 20 at P3.

belief_ratings_control.csv

84 rows, 11 columns. Control was surveyed once, so there is one timepoint (P1) and only the 10 items in blocks A and B.

column description
participant_id 361–467
{item}_P1 the 1–5 response, 10 columns

item_key.csv

20 rows, 10 columns. One row per survey item.

column description
item_id e.g. S2; matches the column prefix in the ratings files
instrument_misinformation_no which of the 10 beliefs it tests (1–9, 11)
block A, B, C or D
question_hi the question as asked, in Hindi
key_d +1 if agreeing is correct, -1 if agreeing is the misinformation
key_meaning key_d in words
fielded_P1_pre, fielded_P1_post, fielded_P2, fielded_P3 1 if asked at that timepoint

Each of the 10 video beliefs has two items, one worded so that agreement is correct and one so that agreement is the misinformation.

response_scale.csv

5 rows. Maps each 1–5 value to its Hindi response option and English gloss.

value meaning
5 strongly believe
4 believe
3 don't know
2 do not believe
1 strongly do not believe

Scoring

Convert a 1–5 response to a belief score using key_d from item_key.csv:

score = (response - 3) * key_d / 2

This gives −1 to +1, where +1 = rejects the misinformation, −1 = believes it, and 0 = "don't know". A participant's score for a block is the mean over that block's items.


qualitative/

qual_responses.csv

350 rows, 12 columns. Interviews conducted individually after the video. Control was not interviewed, so this file covers the two video conditions only.

column description
participant_id 1–360
condition culturally adaptive or culturally neutral
Q1 what the video was about, and your reaction free text
Q2 who the video was made for free text
Q3 had you met these claims before free text
Q4 what a firm believer would say free text
Q5 what felt familiar; made for your community? free text
Q6 how the room took it; disagreement; watching others free text
Q7 would you share it; anyone you would not show free text
Q8 where these topics arise in your life free text
Q9 what the study was about; expected answer? free text
Q13 real person or computer-generated real person, computer-generated, or cannot say

Free text is in Hindi, romanised Hindi, or a mix. Personal names have been replaced with [NAME]. The presenter's name is kept: she is Ganga Devi in the culturally adaptive video and Amy in the culturally neutral one.

demographics.csv

434 rows, 15 columns. All three conditions.

column description
participant_id 1–467
group Culturally adaptive, Culturally neutral, or No-video control
Age band 18 साल से कम, 18-24, 25-34, 35-44, 45-54, 55 - 64, 65+
Owns a mobile phone हाँ / नहीं
Whose phone मेरा खुद का (own), पति (husband), बेटा-बेटी (child), सास/ससुर (in-law), अन्य (other)
Operates phone unaided हाँ, खुद (yes, alone) / थोड़ी मदद से (with help) / नहीं चलाना आता है (cannot)
Phone has internet हाँ / नहीं / पता नहीं (don't know)
Heard of AI हाँ / नहीं
Ever searched health info on phone हाँ, खुद (yes, alone) / हाँ, किसी और से करवाकर (yes, via someone) / नहीं
WhatsApp use, YouTube use, Voice search use, Voice typing use, Facebook/Instagram use, Chatbot use frequency band, e.g. हर दिन 2-4 घंटे (2–4 hours daily), सप्ताह में एक बार (weekly), कभी नहीं (never)

Responses are in Hindi as recorded by the interviewer.


Ethics

IRB approved: Protocol #E-7161 (provider interviews), #E-7912 (field study) at MIT. All participants gave informed consent.

Citation

Anku Rani, Kokil Jaidka, Shruti Sharma, Pragya Mahajan, Manisha Wadhwa, Andrew B. Lippman, Pattie Maes, Paul Pu Liang. Durably Reducing Belief in Women's Health Misinformation Through Culturally Adaptive AI Videos. arXiv:2609.19364 (2026). https://arxiv.org/abs/2609.19364

@misc{rani2026durablyreducingbeliefwomens,
      title={Durably Reducing Belief in Women's Health Misinformation Through Culturally Adaptive AI Videos}, 
      author={Anku Rani and Kokil Jaidka and Shruti Sharma and Pragya Mahajan and Manisha Wadhwa and Andrew B. Lippman and Pattie Maes and Paul Pu Liang},
      year={2026},
      eprint={2609.19364},
      archivePrefix={arXiv},
      primaryClass={cs.HC},
      url={https://arxiv.org/abs/2609.19364}, 
}
Downloads last month
170

Paper for ankurani/healthvoices