emg2qwerty: A Large Dataset with Baselines for Touch Typing using Surface Electromyography
emg2qwerty is the largest public surface electromyography (sEMG) dataset to date, comprising 1,135 sessions from 108 participants performing touch typing on a QWERTY keyboard. The dataset captures wrist-based sEMG signals (32 channels, 2000 Hz sampling rate) synchronized with keystroke ground truth, totaling 346.4 hours of data and 5.26 million keystrokes. Designed to enable keyboard-free text input through decoding of typing intent from neuromuscular activity, the dataset supports research in sequence-to-sequence learning, cross-user generalization, domain adaptation, and neuromotor interfaces for AR/VR and accessibility applications.
AI-generated description, may include mistakesLoading demographics…
Coming soon. Per-file data-quality summaries are precomputed by the NEMAR processing pipeline. The static aggregate is on the way — tracked at nemar-cli#511.
Files
How to use the data (for agentic research) license, citation, download commands
What it is
- Modalities
- EMG
- Participants
- 108
- Size
- 223 GB
- Tasks
- typing
License and terms
- License
- CC-BY-NC-SA-4.0
- Note
- Non-commercial use only (CC-BY-NC-SA-4.0).
- Recommended citation
- Sivakumar, V., Seely, J., Du, A., Bittner, S. R., Berenzweig, A., Bolarinwa, A., Gramfort, A., & Mandel, M. I. (2026). emg2qwerty: A Large Dataset with Baselines for Touch Typing using Surface Electromyography (Version v2.0.0) [Data set]. NEMAR. https://doi.org/10.82901/nemar.nm000104
Where the bytes are
- Latest version (always current)
- https://data.nemar.org/nm000104/latest/
- This version (v2.0.0)
- https://data.nemar.org/nm000104/v2.0.0/
How to download
- The dataset
-
nemar dataset download nm000104Clones and fetches in one step. Content under stimuli/ and derivatives/ is skipped by default because those trees can be large; add --stimuli --derivatives for the whole thing. - A subset, one step
-
nemar dataset download nm000104 --subjects sub-01,02Also filters by --sessions, --tasks, --runs, --datatypes, --include and --exclude. - A subset, step 1
-
nemar dataset clone nm000104Clones git-annex pointers only; fetches no file content. Creates ./nm000104. - A subset, step 2
-
cd nm000104The get command below reads the clone's annex, so it only works from inside the clone. - A subset, step 3
-
nemar dataset get <files>Pulls the files you actually need. Skips stimuli/ and derivatives/ unless the path you ask for is under one of them. - One small file
- https://data.nemar.org/nm000104/v2.0.0/participants.tsv A direct HTTPS fetch works for any single file.
Assess fit without downloading
- Participants table
- https://data.nemar.org/nm000104/v2.0.0/participants.tsv
- Dataset description
- https://data.nemar.org/nm000104/v2.0.0/dataset_description.json
- Directory listing
- https://data.nemar.org/nm000104/v2.0.0/?format=json
- Catalog record
- https://api.nemar.org/datasets/nm000104