Research

What we are testing.

This page is the working record. The question, what would prove us wrong, what we have run so far, and what we still do not know. It is updated as things break. Nothing here is a result yet.

The question

Does contact force have to be measured, or can it be estimated from video well enough that measuring it is a waste of money?

Five funded companies are currently betting that force matters for robot manipulation. None of them have published the experiment that would settle it. Nobody runs it, partly because negative results are hard to fund, and partly because if you sell force data you would rather not find out it does not help.

Why it is unresolved

A hand squeezing hard without moving and a hand at rest occupy nearly the same pose and nearly the same pixels. On that argument, force is not recoverable from kinematics at all. But recent work estimates tactile signal from egocentric video with reasonable accuracy. Both cannot be fully right, and neither has been tested on skilled industrial hand work.

Why we can answer it

The test needs paired data: real force, measured, alongside video, on tasks where force decides the outcome. That means access to workshops where deformable work is still done by hand, and permission to instrument the people doing it. One of our sites is a family brass factory in Moradabad.

Preregistration

v0.1 · written before collection

Hypotheses, thresholds and analysis plan, fixed in advance so the result cannot be chosen after the fact. Revisions are dated and the earlier version stays visible.

H1 Forearm surface EMG predicts fingertip normal force during isometric grip. Pre-set threshold: R2 above 0.70 within a single session, with prediction error under 1.5 N across a 0 to 10 N range. Below 0.50 we treat single-channel EMG as insufficient and add channels before continuing.
H2 The EMG to force mapping degrades when the band is removed and replaced. We expect degradation and want its size. Measured as the change in fitted slope and in R2 across three conditions: first wearing, replaced on a skin mark, and replaced 2 cm off the mark. This number decides whether a wearable needs recalibration every session, which is a hardware requirement, not a software one.
H3 Muscle activation precedes measured force at contact onset. Expected lead of 30 to 100 ms, measured by cross correlation between the EMG envelope and the force trace at discrete contact events. If present, the signal contains intent before the physical consequence, which video cannot contain by construction.
H4 Skilled and unskilled performance of the same task differ in force, not only in kinematics. Same task, same bench, worker with more than ten years in the trade against a first year apprentice. Pre-set comparisons: peak force, cycle-to-cycle variance, and the tightness of the overlaid force traces. Direction predicted in advance: the skilled worker uses less peak force and is more consistent.
H5 Policies trained on video plus force outperform policies trained on video alone, on deformable tasks. The main test, and the expensive one. Identical architecture, identical tasks, identical training budget. The only difference is the force channel. Reported as the difference in success rate with confidence intervals, whichever way it goes.

What would make us stop

If video-based contact estimation predicts our measured force at accuracy comparable to direct measurement on the same footage, then measuring force in the field is not worth its cost. In that case our position changes from "we measure force" to "we provide the ground truth those estimators are validated against," which is a smaller business and a real one. We will say so publicly rather than quietly changing the pitch.

Falsification condition, set in advance

Analysis plan

StepMethodFixed in advance
Force referenceFSR calibrated against a load cell, loading and unloading curves separatereport RMSE in N
EMG conditioningBandpass 20 to 450 Hz, notch at 50 Hz, full-wave rectify, RMS over a 150 ms windowwindow fixed, not tuned
MappingLinear and power-law fits from envelope to newtons, reported togetherno model selection after seeing results
AgreementBland-Altman in addition to R2, because correlation hides biasboth reported
Lead timeCross correlation, peak lag at discrete contact onsetsms, with spread
ExclusionsSessions with failed sync, saturated channels, or a slipped bandcriteria written before collection

Sample: n = 1 for the pilot, which establishes feasibility and nothing about generality. Then five workers, three sessions each, across two trades. We will say which claims rest on which n.

Open questions

Things that could sink this

Each one is live. Where we hold a view, it is marked as a view and not as a finding. We have written to the authors of most of the work below and asked them to tell us we are wrong.

01Is contact force recoverable from egocentric video?

TouchAnything estimates bimanual tactile signal from video. MiTaS reports vision-only success at 31 percent, visual-tactile at 54 percent, and all modalities at 80 percent on contact-rich tasks. The two point in opposite directions, and neither has been evaluated on industrial deformable work.

Our view: estimation likely works where contact is visible and fails where it is not, which is exactly the case in cloth work where the material occludes the fingers. Untested.

02Is shear the binding constraint rather than normal force?

When cloth slips, what fails is friction acting along the surface, not pressure into it. Normal force may be the smaller half of the story for fabric. We do not measure shear in v0 and say so on every page that mentions the rig.

Our view: for brass, normal force dominates. For cloth, shear probably matters more. If that holds, the first release is more useful for metal than for textiles, and we should say which.

03Does online adaptation remove the need for offline human data?

OmniTacTune learns a tactile residual online from base-policy rollouts and explicitly requires no offline tactile demonstrations. It is by the same lead author as MimicTouch, which collected human tactile demonstrations.

Our view: online adaptation still needs something to adapt from, and cannot invent contact statistics for materials it has never touched. Weakly held.

04Does scene diversity beat depth at our budget?

EgoVerse reports that scene diversity strongly affects generalisation under limited data budgets. Our plan is the opposite: few sites, many hours, one trade at a time. Their own live dataset has 240 scenes across roughly 4,000 hours, which is scene-thin by the same measure.

Our view: diversity matters for perception, depth matters for contact statistics of a specific material. Convenient for us, which is a reason to distrust it.

05Is action-space distance the right axis, and where does force sit on it?

The strongest published critique of egocentric collection argues that the real axis is how far the captured data sits from the robot's action space, and that human video is furthest away because five fingers have nowhere to go on a two-finger gripper.

Our view: that is true for kinematics and may be false for force. A joint configuration carries morphology. A scalar newton value at a contact point does not, and both a parallel-jaw gripper and a five-fingered hand have a force command. If so, force is the one part of human demonstration sitting near zero distance. This is a hypothesis, and it is H5 above.

06Can EMG-derived force compose with the tactile representations the field is building?

T3 learns tactile representations across 3M images from 13 sensors. If coarse scalar force cannot compose with that, our data risks being clean, correct and unused.

Our view: unknown. This is the question we most want a researcher to answer for us.

Lab notes

Newest first · failures included

2026-09-22

EMG sensor ordered, protocol fixed before it arrives

Single-channel dry-electrode sensor, chosen over the better-documented alternative because it uses dry contacts on a strap. Gel electrodes give a cleaner signal and tell us nothing about the conditions a wearable will actually face.

Protocol and analysis plan written and frozen before the hardware arrives. Recording at 1000 Hz for EMG and 100 Hz for the force reference, on the same clock.

2026-09-15

Moved from a glove to a forearm band

The glove design puts force sensors on the fingertips. That is the correct place to measure and the wrong place to put anything. A craftsman is paid by the piece. Cover the pad of his finger and he can no longer feel the material, his work slows, and the device comes off.

The sensing problem was never the hard part. The constraint is that the instrument has to disappear. Everything after this decision is downstream of that, including a harder sensing problem: the band infers force from muscle activity rather than measuring it at the contact.

Force sensors are not abandoned. They become the calibration reference, worn for short sessions rather than a full shift. There is no way to turn muscle signal into newtons without measuring real force alongside it.

2026-09-08

Mapped the field properly and found we are not first

82 companies in robot training data. Five already working on force. Three building tactile gloves, two building EMG bands, one of those funded at roughly 73 million dollars within five months of starting.

This changed the positioning rather than the plan. Sensing is not a defensible wedge. What is left is the task distribution, skilled trades where the material deforms, and access to the floors where that work still happens.

2026-09-01

Wrote to the researchers before spending money

Contacted authors behind the main papers in this area, at Georgia Tech, ETH Zurich, MIT and IISc, with one question each and a request to tell us the thesis is wrong.

The specific questions: is force or shear the missing modality, does online adaptation supersede offline collection, and does the scene diversity finding invalidate a depth-first collection plan.

Entries are added when something changes, not on a schedule.

Method

Two devices on one clock. Video says what the hands did. The band says how hard they pressed. Neither is useful without the other being time-aligned to it.

01

Capture

Head camera at 1080p30. Forearm band, 8 EMG channels at 1000 Hz. Both write locally.

02

Sync

LED on the band flashes three times at session start and end. Detected in video, mapped to frame index, drift corrected across the session.

03

Condition

Bandpass, notch, rectify, RMS envelope. Force reference converted to newtons using the per-unit calibration curve.

04

QC

Automated checks, then human review. A session either passes or is discarded with a recorded reason.

05

Release

Packaged in the format labs already load, with the datasheet and the limitations attached.

What passes QC

Every check is automated except the last. A failed session is kept with its failure reason rather than deleted, because the failure distribution is itself informative.

CheckPasses ifWhy it exists
Sync detectedFlash found at both ends, mapped lag under one frameAn unsynced session is unusable at any quality
Clock driftCorrected residual under 20 ms across the sessionTwo clocks diverge over hours
Calibration ageForce reference calibrated within 7 daysSensors drift under repeated load
Channel saturationUnder 1 percent of samples at railA clipped channel silently reports the wrong force
Dropped framesUnder 0.5 percentGaps break the alignment we just established
Hand visibilityAbove a stated threshold, computed the same way as published datasetsComparability
Band placementPhoto taken at donning, position loggedH2 says placement moves the mapping
ConsentSigned, versioned, in the worker's language, revocableNon-negotiable
Human reviewOne reviewer confirms the task matches the labelAutomated checks do not catch a mislabelled task

Data card

Structure fixed · nothing collected yet

Following the datasheets-for-datasets convention. Published with the first release and versioned alongside it, including the parts that make the data less useful.

FieldPlanned content
MotivationTesting whether contact force improves manipulation learning on deformable materials
CompositionPaired egocentric video, forearm EMG, derived force estimate, hand orientation, task and material labels
CollectionReal production work in operating workshops, not staged tasks. Workers paid above local shift rate.
Skill annotationYears in trade per worker. Same task recorded across experience levels at the same bench.
PreprocessingStated filter parameters, calibration curves shipped per unit, raw counts retained alongside newtons
Not measuredShear. Tactile texture. Absolute per-finger force with certainty, since the band infers rather than measures.
Known limitationsTwo trades, two geographies, small worker count in v0. Sensor variance and hysteresis quantified and published.
Consent and revocationVersioned consent, revocable after the fact, with the process for removing a worker's sessions
LicenceOpen. First release free to use, including commercially.

People

Consent is written, versioned, read aloud in the worker's language, and revocable after recording. Workers are paid roughly double the local rate for the hours they wear the equipment, and the rate does not depend on what the data turns out to be worth. A physical stop control ends recording immediately, without a phone and without asking anyone.

Why this section exists

This work instruments people at their place of employment, where slowing them down costs them money. Treating that as a procurement detail rather than a research ethics question would be a mistake, and any lab using the data will want to know how it was obtained.

How we work

We use language models heavily and it is worth being specific about where they help and where they do not, since the failure modes matter more than the wins.

Reading

Annotated · what it claims, what it costs us

The strongest argument against what we are doing. Proposes action-space distance as the real axis and places egocentric human data furthest from it. Read this before anything supportive.

Bimanual tactile estimation from egocentric video. If this holds at production accuracy on our task distribution, the band is unnecessary and our position becomes ground-truth provider.

MiTaS

2026

The counterweight, with numbers. Vision-only 31 percent, plus tactile 54 percent, all modalities 80 percent across five contact-rich tasks. Robot-side sensing, not human-side.

The spec we build against, and the finding that challenges our site plan. Also the sentence this company rests on: video alone cannot capture the contact signals manipulation depends on.

EgoMimic

Georgia Tech · 2024

Why the category exists. An additional hour of human hand data beats an additional hour of robot data, with large reported gains from co-training.

MimicTouch

Georgia Tech · CoRL 2024

Human tactile demonstrations collected directly from human hands, with an audio channel alongside. Someone solved a version of our hardware problem already.

Online tactile residual learning that explicitly needs no offline tactile demonstrations. Same lead author as MimicTouch, which makes it the most interesting tension in the field.

PolyTouch

MIT · ICRA 2025

Camera-based tactile plus an acoustic channel that hears contact beginning and breaking, which is precisely where resistive force sensing is weakest. Designed for durability, which is our constraint too.

T3

MIT · CoRL 2024

Tactile representations across 3M images and 13 sensors. Determines whether scalar force data can compose with what the field is building, or is stranded.

DROID

RSS 2024

The only real precedent for distributed collection across sites you do not control. Read for the quality problems that only surfaced after collection had finished.

LeRobot

Hugging Face · ICLR 2026

The format labs actually load. We build to it from the first session, because converting later costs adoption.

Status

Kept current, including when there is nothing to report.

ItemState
Hours collected0
Preregistrationv0.1 published, before collection
Force referencesensors in hand, calibration in progress
Bandsingle-channel pilot, hardware ordered
Sitestwo workshops in conversation
Fundingself-funded
First releasenot scheduled, will be open when it happens