Fylgja Open the platform

Hosted ethology platform

Run the study,
not the pipeline.

Fylgja turns raw behavioural video into tracked data, with a validation report attached to every run. Pose estimation runs on hosted GPUs, the analysis contract is documented, and each export carries its provenance.

A joint validation manuscript with Sahlgrenska neuroscience researchers is in preparation.

  1. Sourcecamera files
  2. Preprocesssplit · mask · enhance
  3. PoseGPU inference
  4. Metricsdocumented contract
  5. Reviewhuman where it counts
  6. Exportbundle + provenance

27

keypoints tracked per animal, from the SuperAnimal-TopViewMouse model (HRNet-W32)

2.59 px

pose RMSE with 89.7% mAP and 91.0% mAR after fine-tuning, measured on a 21-image held-out split

No install

pose estimation and fine-tuning run on hosted GPUs while analysis, review and exports live in the browser; no CUDA and no local environment to maintain

1–3 h

what one published study reports for labelling five behaviours in a ten-minute mouse recording

Where we are

Stated plainly.

Platform

Built end to end and running in development: ingest, GPU inference, analysis, review, validation, export.

Partner study

Joint manuscript with Sahlgrenska neuroscience researchers in preparation, under a protocol that blind-compares Fylgja against human consensus and the lab's existing workflow.

Targets, not claims

The validation protocol targets reproducing the blinded human conclusion in roughly 96% of experiments while cutting analysis labour by around 87%. Those numbers are reported only once the blinded validation programme earns them.

Vision

The infrastructure ethology never had.

Genomics has reference genomes and shared pipelines. Imaging has standard formats and shared compute. Behaviour research still depends on one lab's scripts and one student's laptop, and that dependency quietly sets the ceiling on what the field can ask.

Fylgja is working toward the first hosted ethology platform: the whole measurement chain, from raw video to a validated result, available in a browser and metered by the research minute. No cluster to rent, no environment to maintain, and no pipeline to keep running.

That means two things. A lab can ask a behavioural question well without first building infrastructure. And every review, correction, and accepted result is designed to feed a cross-lab corpus with full provenance, which is what a canonical behavioural model needs to exist at all.

The lower the bar to rigorous measurement, the more science gets done.

The bottleneck

Behaviour research still runs on bespoke tooling.

The most common analysis method remains a researcher watching video with a stopwatch. One peer-reviewed study reports one to three hours of expert annotation for every ten minutes of footage. Researchers have also reported around twenty-five person-hours to annotate a single hour of mouse video (2010).

The machine-learning path excludes most of the field. Pose estimation asks for GPUs, CUDA, Python, and often 50 to 200 hand-labelled frames before it answers anything.

So labs without an engineer score by hand, and much of their video stays unquantified.

Human scoring carries its own variance. In one published benchmark, two trained observers agreed on simple left/right turns at frame-wise κ ≈ 0.64–0.65, and published inter-observer reliability across behaviours spans roughly κ 0.64–0.91. Against that background, trust matters more than convenience. An analysis bug does not inconvenience a user; it invalidates a figure.

Methodology

The analysis contract, in the open.

This is the reference pipeline from the three-chamber study Fylgja was built on, in four phases from camera file to validated result. The numbers below are that study's; a project can override them, and the values a run actually used are recorded with the run. Every stage leaves artifacts you can inspect.

Phase 0

Preprocessing

Segmentation
pretest, locomotion, stranger
Spatial split
left and right arena cameras
Masking
chamber polygons, gates, middle arena
Enhancement
brightness +50, H.264
Rates
30 fps source, analysed at 10 fps
Output
320 × 480 px per camera

Phase 1

Pose estimation

Model
SuperAnimal-TopViewMouse, HRNet-W32
Keypoints
27: head 8, body 6, tail 7, limbs 6
Batching
16 frames, detector batch 8
Adaptation
available per video
Hardware
hosted GPUs, no local install
Output
H5, CSV, JSON

Phase 2

Fine-tuning

Frames
409 curated from 110 recordings
Split
388 train, 21 validation
Schedule
AdamW, 1×10−4, crop 448
Epochs
200 max, early stop at 66
Augmentation
affine, Gaussian noise, motion blur
Result
89.69% mAP, 90.95% mAR, 2.59 px RMSE
Training progression, same validation set
EpochmAPmARRMSE
1071.26%73.33%4.17 px
3087.88%89.05%2.80 px
6089.69%90.95%2.59 px

Phase 3

Behaviour metrics

Chamber occupancy
point-in-polygon, 3-frame hysteresis; no direct top↔bottom transitions
Sniffing
nose 1.5–12 px/frame with body velocity below 6 px/frame, ≥5 frames; labelled by cup proximity or exploratory
Grooming
body elongation below 2.5 and velocity below 2 px/frame for ≥5 frames
Stranger preference
time within 50 px of the stranger cup against the empty cup
Distance
centre-path length, confidence filtered
Validation
cross-checked against an independent centroid tracker

Quality gate

Tracking agreement between the pose pipeline and an independent centroid tracker sorts each recording: under 15 px excellent, 15–30 px good, 30–50 px moderate, over 50 px poor.

Validation as an artifact

Validation is a step of the pipeline, not a hidden check. Its report is written on every meaningful change and travels with the exported research bundle.

Platform

The model is one part of the job.

Organisation, review, and evidence are the actual work of a study, so the platform is built around those.

  • Lab-scoped workspaces

    Collaboration and authorisation happen at the lab boundary; projects hold the work.

  • Sources and setup

    Resumable uploads, timestamp segments, and apparatus geometry captured as explicit inputs.

  • Annotation

    Review tasks are claimed, completed, and published. Overrides stay versioned, never hidden.

  • Jobs

    Preprocessing, inference, analysis, validation and export run as tracked jobs; the GPU work goes to hosted hardware.

  • Analysis and statistics

    Versioned configs, run diffs, cohort statistics, individual trajectories, robustness checks.

  • Validation and export

    The run is validated, then packaged as an immutable research bundle with its provenance.

Every result must answer three questions.

What ran
the exact pipeline version and configuration, recorded per run
What changed
human overrides and configuration edits, kept as inputs
What produced this
artifacts, provenance, and lineage behind every number

Bring one recording.

Create a project, upload a video, and watch the pipeline turn it into tracked, reviewed data with its validation report attached.

Open the platform