SIGN-LAB
Teaching machines to understand human motion.
From motion to meaning.
An inside view of recorded perception, model inputs, and early translation.
- Real human motion
- Real model inputs
- Recorded benchmark output

AI learned to read. Then speak. Then see. Understanding human movement asks something different. A handshape alone is not language. A facial expression alone is not language. Meaning depends on how signals unfold together through time.
UMI is researching the layer between human movement and machine understanding. bitsign is where we begin.
This is a research instrument you can explore — a recorded benchmark inspection, not a live inference service or a claim of a solved translator.
One signer. Multiple information streams.
Detection produces estimated reference points, not verified physical ground truth. The counts below are the observed representation recorded during this inspection. A count of available landmarks is not a count of points successfully tracked in every frame.
- 0
- Face points per detection
- 0
- Points per detected hand
- 0
- Body points per pose
Observed representation
Observed representation
Observed representation
Geometry. Appearance. Motion.
There are 21 points per detected hand in the recorded output. Geometry describes the wrist, joints, and fingertips.
Visual hand crops retain image information a skeleton does not express. The hand skeleton itself is not the whole translation input.
Where real coordinate traces are shown, they break across missing observations. We do not interpolate failures into successful-looking tracking.
The detector observes more than the current representation retains.
Recorded face detections contain 478 points per face. The current downstream face image primarily preserves eye and mouth regions against a neutral background; it does not consume the full facial mesh as a complete facial representation.
Facial information is part of the research question. Observing a signal and preserving its linguistic contribution are different problems. We do not claim eyebrow grammar, emotion, identity, or nonmanual meaning is reliably decoded by this model.
Movement has a frame of reference.
The recorded detector has 33 body points. The current downstream body representation uses 7selected points, with two spatial coordinates each: nose, two shoulders, two elbows, and two wrists. These selected points are not all of the body's movement and do not exhaust the translator's input.
- Nose
- Left shoulder
- Right shoulder
- Left elbow
- Right elbow
- Left wrist
- Right wrist
One frame is a state. Meaning unfolds across a sequence.
The benchmark footage is 30 frames per second. Clip 01 has 431 frames over approximately 14.37 seconds. These are media properties, not inference speed. The research objective is connecting changing hands, face, posture, and context through time.
- 0.0s
0 - 1.3s
39 - 2.6s
78 - 3.9s
118 - 5.2s
157 - 6.5s
196 - 7.8s
235 - 9.1s
274 - 10.4s
313 - 11.8s
353 - 13.1s
392 - 14.4s
431
Real frames sampled from Clip 01 at 30 fps · media properties, not inference timing
Conceptual flow
- 01→
Video
- 02→
Perception
- 03→
Visual + spatial signals
- 04→
Temporal context
- 05
Early language output
This diagram uses image crops and real recorded body geometry, not fictional vector values or activation heatmaps. The current system still uses appearance-bearing visual crops; appearance-free motion understanding is a research direction, not a deployed privacy property.
Feature Clip 01
Early model translation
“Alkaline chemicals change the color of acids and basics.”
Reference meaning
“The cabbage juice changes color depending on how acidic or basic (alkaline) the chemical is.”
MEANING MATCH · 71%
Semantic similarity against reference
Similarity score × 100; not translation accuracy
The output retains related words about chemicals, color, acids and bases. It omits cabbage juice and changes the relationships between the concepts. Topic overlap and a faithful translation are different achievements.
This is a text comparison — sentence-embedding cosine similarity — not evidence about what the model internally understood or which perception event caused the error. It does not show 71% of meaning understood, translation correctness, or base-model accuracy across tasks. Only Clip 01 currently has this measured similarity evidence.
Recorded benchmark inspection.
9 clips, 3,045 frames, 101.5 seconds total, all 30 fps. These numbers describe this small inspection set, not total training scale or generalization quality. The full model view always preserves its six windows. Missing media shows an honest awaiting state — never a substituted sample.
- Duration
- 14.37s
- Frames
- 431
- Rate
- 30 fps
- Final index
- 430
Select clip
Clip 01 coverage
Percentage of frames in which each stream selected a detection. Coverage is not tracking accuracy or proof of correct handedness.
- Left hand selected
- 76.1%
- Right hand selected
- 91.6%
- Face detected
- 100%
- Pose detected
- 100%
Recorded output
Early model translation
Alkaline chemicals change the color of acids and basics.
Reference meaning
The cabbage juice changes color depending on how acidic or basic (alkaline) the chemical is.
MEANING MATCH · 71%
Semantic similarity against reference
Similarity score × 100; not translation accuracy
FLEURS-ASL / Google · CC BY-SA 4.0. Sentence segment; perception overlays and public model-input visualization added. No translation rerun.
We don't hide the failure. We instrument it.
A missing selection. A held input. An incorrect relationship in a sentence. Each gives us a question we can measure and test.
Observed input events
- Clip 01 begins with 69 frames (2.3 seconds) without a selected left-hand input.
- Across the inspection there are six frames with a duplicate assignment to the two hand streams. We present only the observable duplicate state.
- Crops can be blank or reuse previous input when a fresh crop is unavailable. A hand may be outside the frame or occluded — a missing selection alone is not a confirmed detector error.
Coverage means a stream selected a detection. It is not tracking accuracy or proof of anatomically correct handedness. No claim that these events caused a translation error has been established.
The learning loop
- 1Observe
- 2Measure
- 3Hypothesize
- 4Test
- 5Benchmark
- 6Learn
Retain baseline evidence, isolate changes, compare outcomes, promote measured improvements. A change that looks better is a hypothesis. A measured improvement is evidence.

