UMI
bitsign.ai LabFrom motion to meaning.
Back to bitsign.ai
01Human motion

SIGN-LAB

Teaching machines to understand human motion.

From motion to meaning.

An inside view of recorded perception, model inputs, and early translation.

  • Real human motion
  • Real model inputs
  • Recorded benchmark output
A person signing, captured against a quiet studio background.
Recorded benchmark inspection

AI learned to read. Then speak. Then see. Understanding human movement asks something different. A handshape alone is not language. A facial expression alone is not language. Meaning depends on how signals unfold together through time.

UMI is researching the layer between human movement and machine understanding. bitsign is where we begin.

This is a research instrument you can explore — a recorded benchmark inspection, not a live inference service or a claim of a solved translator.

02Perception

One signer. Multiple information streams.

Detection produces estimated reference points, not verified physical ground truth. The counts below are the observed representation recorded during this inspection. A count of available landmarks is not a count of points successfully tracked in every frame.

0
Face points per detection

Observed representation

0
Points per detected hand

Observed representation

0
Body points per pose

Observed representation

03Hands

Geometry. Appearance. Motion.

There are 21 points per detected hand in the recorded output. Geometry describes the wrist, joints, and fingertips.

Visual hand crops retain image information a skeleton does not express. The hand skeleton itself is not the whole translation input.

Where real coordinate traces are shown, they break across missing observations. We do not interpolate failures into successful-looking tracking.

04Face

The detector observes more than the current representation retains.

Recorded face detections contain 478 points per face. The current downstream face image primarily preserves eye and mouth regions against a neutral background; it does not consume the full facial mesh as a complete facial representation.

Facial information is part of the research question. Observing a signal and preserving its linguistic contribution are different problems. We do not claim eyebrow grammar, emotion, identity, or nonmanual meaning is reliably decoded by this model.

FACE MESH
FACE INPUT
05Body

Movement has a frame of reference.

The recorded detector has 33 body points. The current downstream body representation uses 7selected points, with two spatial coordinates each: nose, two shoulders, two elbows, and two wrists. These selected points are not all of the body's movement and do not exhaust the translator's input.

  • Nose
  • Left shoulder
  • Right shoulder
  • Left elbow
  • Right elbow
  • Left wrist
  • Right wrist
06Time

One frame is a state. Meaning unfolds across a sequence.

The benchmark footage is 30 frames per second. Clip 01 has 431 frames over approximately 14.37 seconds. These are media properties, not inference speed. The research objective is connecting changing hands, face, posture, and context through time.

  • Clip 01 at frame 0, 0.0 seconds0
    0.0s
  • Clip 01 at frame 39, 1.3 seconds39
    1.3s
  • Clip 01 at frame 78, 2.6 seconds78
    2.6s
  • Clip 01 at frame 118, 3.9 seconds118
    3.9s
  • Clip 01 at frame 157, 5.2 seconds157
    5.2s
  • Clip 01 at frame 196, 6.5 seconds196
    6.5s
  • Clip 01 at frame 235, 7.8 seconds235
    7.8s
  • Clip 01 at frame 274, 9.1 seconds274
    9.1s
  • Clip 01 at frame 313, 10.4 seconds313
    10.4s
  • Clip 01 at frame 353, 11.8 seconds353
    11.8s
  • Clip 01 at frame 392, 13.1 seconds392
    13.1s
  • Clip 01 at frame 431, 14.4 seconds431
    14.4s

Real frames sampled from Clip 01 at 30 fps · media properties, not inference timing

07Representation

Conceptual flow

  1. 01

    Video

  2. 02

    Perception

  3. 03

    Visual + spatial signals

  4. 04

    Temporal context

  5. 05

    Early language output

This diagram uses image crops and real recorded body geometry, not fictional vector values or activation heatmaps. The current system still uses appearance-bearing visual crops; appearance-free motion understanding is a research direction, not a deployed privacy property.

08Meaning

Feature Clip 01

Early model translation

“Alkaline chemicals change the color of acids and basics.”

Reference meaning

“The cabbage juice changes color depending on how acidic or basic (alkaline) the chemical is.”

MEANING MATCH · 71%

Semantic similarity against reference

Similarity score × 100; not translation accuracy

The output retains related words about chemicals, color, acids and bases. It omits cabbage juice and changes the relationships between the concepts. Topic overlap and a faithful translation are different achievements.

This is a text comparison — sentence-embedding cosine similarity — not evidence about what the model internally understood or which perception event caused the error. It does not show 71% of meaning understood, translation correctness, or base-model accuracy across tasks. Only Clip 01 currently has this measured similarity evidence.

09The lab

Recorded benchmark inspection.

9 clips, 3,045 frames, 101.5 seconds total, all 30 fps. These numbers describe this small inspection set, not total training scale or generalization quality. The full model view always preserves its six windows. Missing media shows an honest awaiting state — never a substituted sample.

Duration
14.37s
Frames
431
Rate
30 fps
Final index
430

Select clip

Clip 01 coverage

Percentage of frames in which each stream selected a detection. Coverage is not tracking accuracy or proof of correct handedness.

Left hand selected
76.1%
Right hand selected
91.6%
Face detected
100%
Pose detected
100%

Recorded output

Early model translation

Alkaline chemicals change the color of acids and basics.

Reference meaning

The cabbage juice changes color depending on how acidic or basic (alkaline) the chemical is.

MEANING MATCH · 71%

Semantic similarity against reference

Similarity score × 100; not translation accuracy

FLEURS-ASL / Google · CC BY-SA 4.0. Sentence segment; perception overlays and public model-input visualization added. No translation rerun.

10What we are learning

We don't hide the failure. We instrument it.

A missing selection. A held input. An incorrect relationship in a sentence. Each gives us a question we can measure and test.

Observed input events

  • Clip 01 begins with 69 frames (2.3 seconds) without a selected left-hand input.
  • Across the inspection there are six frames with a duplicate assignment to the two hand streams. We present only the observable duplicate state.
  • Crops can be blank or reuse previous input when a fresh crop is unavailable. A hand may be outside the frame or occluded — a missing selection alone is not a confirmed detector error.

Coverage means a stream selected a detection. It is not tracking accuracy or proof of anatomically correct handedness. No claim that these events caused a translation error has been established.

The learning loop

  1. 1Observe
  2. 2Measure
  3. 3Hypothesize
  4. 4Test
  5. 5Benchmark
  6. 6Learn

Retain baseline evidence, isolate changes, compare outcomes, promote measured improvements. A change that looks better is a hypothesis. A measured improvement is evidence.

Motion is information.

From motion to meaning.

Explore the lab