The 2026 PNPL Competition: Word Classification and Efficient Cross-Subject Generalisation in LibriBrain100
Abstract
The 2026 PNPL competition introduces cross-subject and within-subject word classification tracks using an expanded MEG dataset to advance non-invasive speech-decoding brain-computer interfaces toward clinical feasibility.
The ambition of the 2025 PNPL competition (Landau et al., 2025) was to launch a multi-year curriculum for non-invasive speech decoding. Designed to progress from foundational tasks toward the linguistic complexity required for a practical brain-computer interface (BCI), it set the stage with speech detection and phoneme classification tasks. Winning submissions reached F1-macro scores of 95.6% and 73.6% on the respective tasks (Elvers et al., 2026), highly significant advances. This success was built on the LibriBrain dataset (Özdogan et al., 2025), the largest within-subject MEG dataset recorded at the time with {sim}50 hours of data for one subject. However, while within-subject scale drives strong decoding performance, a practical BCI must generalise to new users from minutes of data, not hours. The 2026 PNPL competition responds to this challenge with LibriBrain100 (Mantegna et al., 2026), an extended LibriBrain dataset with 32 additional subjects ({sim}40 minutes each) plus even more within-subject data ({sim}80 hours). Advancing the curriculum of tasks to focus on word classification, two complementary tracks are presented in this competition: the Deep track targets within-subject word classification at scale, aiming at the best possible performance; the Broad track targets cross-subject generalisation, progressively reducing the amount of subject-specific fine-tuning data from {sim}40 to {sim}20 to {sim}10 minutes, a duration that falls within a clinically feasible range and brings us a step closer to a non-invasive BCI capable of restoring communication to people living with profound paralysis.
Community
We’re excited to share the paper for the 2026 PNPL Competition on non-invasive speech decoding 🧠
This year we move to word classification from MEG, with two complementary challenges:
• Deep: push decoding performance using ~80 hours of data from a single subject
• Broad: generalise across subjects with progressively less subject-specific data, including fully zero-shot subjects
The competition is open to everyone until 15 October, with $5,000 in prizes. We’ve also released LibriBrain100, reference models, PyTorch data loaders (pip install pnpl), Colab tutorials, and live leaderboards to make it easy to get started.
We would love to see people from the broader ML community take a shot at it — no neuroscience background required!
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- LibriBrain100: One Hundred Hours of Broad and Deep MEG Data for Neural Speech Decoding at Scale (2026)
- A Common Measure of Communication for Speech Brain-Computer Interfaces (2026)
- The First ChineseBabyLM Challenge: training data-efficient and cognitively plausible language models for Chinese (2026)
- Phoneme- vs. Character-Level Targets and Selective State-Space Models for Intracortical Brain-to-Text (2026)
- The Capacity of Thought: Benchmarking Llama 3.2 in Semantic fMRI Neural Language Decoding and Improving the Huth Encoding-Model Baseline (2026)
- Robust Summarization of Doctor-Patient Conversations: TalTech Systems for the Beyond Transcription Challenge (2026)
- Summary of the ChinaVoices Challenge 2026: Data, Tasks, Baseline, and Methods (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.03231 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper