A Foundation Model for Wearable Movement
Data in Mental Health Research

1Carle Illinois College of Medicine, UIUC  ·  2Georgia Institute of Technology  ·  3NIH/NIDDK  ·  4Dartmouth College  ·  5Geisel School of Medicine
*Equal contribution
Abstract

An open foundation model for actigraphy

Wearable movement data is collected by nearly all commercially available smartwatches and is a valuable resource for mental health research, reflecting fine-grained temporal behavioral trends. Despite its promise, the development of foundation models for health wearable modeling remains limited when compared to clinical image and text analysis.

We designed transformers with patch embeddings and used self-supervised masked autoencoder pretraining on minute-level, week-long actigraphy sequences to develop and investigate the Pretrained Actigraphy Transformer (PAT) — an open-source foundation model that combines week-long temporal modeling, psychiatric outcome evaluation, and full reproducibility on public data.

Pretrained on 21,538 U.S. participants from NHANES, PAT-L outperformed or matched the strongest non-foundation baseline on every task we evaluated — benzodiazepine and SSRI use, depression, and sleep abnormalities — with the largest gains on medication-use prediction. Beyond predictive accuracy, PAT provides interpretable attention maps highlighting periods of daily activity most important for each prediction.

The Challenge

Why actigraphy needs a foundation model

Actigraphy has been used in clinical research since the 1970s, but the analytical tooling has not kept pace with what we now have for clinical text and images. Three main gaps motivated PAT.

01

Rich signal, narrow tools

Minute-level actigraphy captures circadian fragmentation, psychomotor slowing, and day-to-day rhythm variability. Classical feature-engineering pipelines remain transparent but rely on handcrafted, task-specific summaries that are sensitive to cohort and device differences.

02

Models built for shorter horizons

CNNs have limited receptive fields; RNNs degrade over long windows. Most wearable deep learning has focused on second-to-minute activity recognition, leaving the multi-day rhythms and circadian structure that emerge only at minute-level resolution over week-long windows largely unmodeled.

03

The need for an open, week-scale model

Existing wearable foundation models tend to fall short on at least one of four fronts: (a) short temporal windows that miss multi-day behavior; (b) proprietary datasets researchers cannot access; (c) released code but withheld pretrained weights; (d) evaluation on activity recognition or general biomarkers rather than psychiatric and sleep outcomes. PAT is positioned to fill this gap.

Method

PAT: patch-based transformer with masked autoencoder pretraining

PAT applies a Vision-Transformer-style patching scheme (ViT; Dosovitskiy et al., 2020) to week-long, minute-level actigraphy (10,080 timesteps). The sequence is divided into non-overlapping patches, each mapped to an embedding via a trainable linear projection (or 1D convolution in conv variants) and combined with fixed sine-cosine positional embeddings.

PAT pretraining and fine-tuning architecture diagram
Figure 2 from the paper: masked autoencoder pretraining (left) and downstream fine-tuning (right).

Pretraining

  • Masked autoencoder pretraining (MAE; He et al., 2022) on 21,538 participants from NHANES — the CDC's public health survey with worn-accelerometer recordings across a demographically diverse U.S. sample (217M+ minute-level data points)
  • 90% masking ratio — encoder sees only 10% of patches, decoder reconstructs the full week
  • MSE loss on all patches — outperformed masked-only loss substantially (0.773 vs 0.541 avg AUC)
  • No smoothing — raw signal preserved fine-grained activity bursts

Fine-tuning

  • Patch embedding + pretrained encoder + task-specific classifier head
  • Released variants: PAT-S (287K), PAT-M (1.00M), PAT-L (1.99M) params (plus PAT Conv-S/M/L variants with 1D-convolutional patch embeddings)
  • Single-GPU fine-tuning — works on free Google Colab
Headline Results

PAT-L outperforms or matches the strongest baseline on every task

Largest gains on medication-use prediction, where psychotropic effects on sleep–wake rhythm and activity levels create actigraphy signatures the model can capture.

+14.8%
AUC vs ConvLSTM
Benzodiazepine prediction
+5.0%
AUC vs CNN-3D
Sleep abnormality detection
5 / 5
tasks improved
vs strongest non-foundation baseline
21,538
NHANES participants
217M+ data points pretraining
Downstream task

AUC  ↑ higher is better
AUC (Area Under the ROC Curve) measures binary-classification quality on a 0–1 scale, where 1.0 is perfect prediction. Bars show average AUC across training-set sizes (n = 500, 1k, 2.5k, full); PAT-L is compared against the strongest non-foundation baselines from Tables VI–VII of the paper.
Interpretability

Attention organizes around daily behavioral rhythms

Unlike most deep learning models for wearable data, PAT provides built-in interpretability through attention weights — no SHAP or LIME required. Population-level attention from the benzodiazepine model shows repeated off-diagonal bands at 24-hour offsets and consistent peaks concentrated around the morning activity transition.

Quantitative confirmation: 24-hour-lag autocorrelation of the attention profile was consistently positive across participants (mean = 0.36, p < 0.0001), and attention was significantly enriched in a ±60-minute window around inferred wake transitions (mean enrichment ratio = 1.29, Wilcoxon p < 0.0001).

Population-level attention organization in PAT-L showing daily rhythms
Figure 5: Population-level attention organization in PAT-L. (a) Attention matrix showing 24-hour off-diagonal bands. (b) Weekly attention profile with daily peaks. (c) 24-hour overlay — peak attention at the morning activity transition.
Deployment

Built for accessibility, not just accuracy

PAT was designed to be useful to the health researcher, not just benchmarked on a leaderboard. Every variant fits on a single consumer GPU and the full stack is open.

⚡

Single-GPU fine-tuning

Every PAT variant fine-tunes on a free Google Colab GPU. PAT-M completes inference on a week-long sequence in ~35 ms — faster than LSTM (117 ms) and ConvLSTM (159 ms) baselines.

🪶

Lightweight by design

287K to 1.99M parameters across variants — PAT-S and PAT-M are comparable to or smaller than the ConvLSTM (1.76M) and 3D-CNN (790K) baselines they outperform.

🔓

Fully open

Pretrained weights, training code, and the underlying NHANES dataset are all publicly available. Tutorial notebooks cover fine-tuning, explainability, and MAE pretraining.

🩺

Privacy-preserving

Local fine-tuning means sensitive wearable data does not need to leave the institution where it was collected — simplifying privacy-preserving workflows.

Cite

BibTeX

@article{ruan2026pat,
  title   = {A Foundation Model for Wearable Movement Data in Mental Health Research},
  author  = {Ruan, Franklin Y. and Zhang, Aiwei I. and Oh, Jenny Y. and Jin, SouYoung and Jacobson, Nicholas C.},
  journal = {IEEE Journal of Biomedical and Health Informatics (J-BHI)},
  year    = {2026},
  url     = {https://ieeexplore.ieee.org/stamp/stamp.jsp?tp=&arnumber=11524051}
}