Research
This page is the working record. The question, what would prove us wrong, what we have run so far, and what we still do not know. It is updated as things break. Nothing here is a result yet.
The question
Does contact force have to be measured, or can it be estimated from video well enough that measuring it is a waste of money?
Five funded companies are currently betting that force matters for robot manipulation. None of them have published the experiment that would settle it. Nobody runs it, partly because negative results are hard to fund, and partly because if you sell force data you would rather not find out it does not help.
A hand squeezing hard without moving and a hand at rest occupy nearly the same pose and nearly the same pixels. On that argument, force is not recoverable from kinematics at all. But recent work estimates tactile signal from egocentric video with reasonable accuracy. Both cannot be fully right, and neither has been tested on skilled industrial hand work.
The test needs paired data: real force, measured, alongside video, on tasks where force decides the outcome. That means access to workshops where deformable work is still done by hand, and permission to instrument the people doing it. One of our sites is a family brass factory in Moradabad.
v0.1 · written before collection
Hypotheses, thresholds and analysis plan, fixed in advance so the result cannot be chosen after the fact. Revisions are dated and the earlier version stays visible.
If video-based contact estimation predicts our measured force at accuracy comparable to direct measurement on the same footage, then measuring force in the field is not worth its cost. In that case our position changes from "we measure force" to "we provide the ground truth those estimators are validated against," which is a smaller business and a real one. We will say so publicly rather than quietly changing the pitch.
Falsification condition, set in advance| Step | Method | Fixed in advance |
|---|---|---|
| Force reference | FSR calibrated against a load cell, loading and unloading curves separate | report RMSE in N |
| EMG conditioning | Bandpass 20 to 450 Hz, notch at 50 Hz, full-wave rectify, RMS over a 150 ms window | window fixed, not tuned |
| Mapping | Linear and power-law fits from envelope to newtons, reported together | no model selection after seeing results |
| Agreement | Bland-Altman in addition to R2, because correlation hides bias | both reported |
| Lead time | Cross correlation, peak lag at discrete contact onsets | ms, with spread |
| Exclusions | Sessions with failed sync, saturated channels, or a slipped band | criteria written before collection |
Sample: n = 1 for the pilot, which establishes feasibility and nothing about generality. Then five workers, three sessions each, across two trades. We will say which claims rest on which n.
Things that could sink this
Each one is live. Where we hold a view, it is marked as a view and not as a finding. We have written to the authors of most of the work below and asked them to tell us we are wrong.
TouchAnything estimates bimanual tactile signal from video. MiTaS reports vision-only success at 31 percent, visual-tactile at 54 percent, and all modalities at 80 percent on contact-rich tasks. The two point in opposite directions, and neither has been evaluated on industrial deformable work.
Our view: estimation likely works where contact is visible and fails where it is not, which is exactly the case in cloth work where the material occludes the fingers. Untested.
When cloth slips, what fails is friction acting along the surface, not pressure into it. Normal force may be the smaller half of the story for fabric. We do not measure shear in v0 and say so on every page that mentions the rig.
Our view: for brass, normal force dominates. For cloth, shear probably matters more. If that holds, the first release is more useful for metal than for textiles, and we should say which.
OmniTacTune learns a tactile residual online from base-policy rollouts and explicitly requires no offline tactile demonstrations. It is by the same lead author as MimicTouch, which collected human tactile demonstrations.
Our view: online adaptation still needs something to adapt from, and cannot invent contact statistics for materials it has never touched. Weakly held.
EgoVerse reports that scene diversity strongly affects generalisation under limited data budgets. Our plan is the opposite: few sites, many hours, one trade at a time. Their own live dataset has 240 scenes across roughly 4,000 hours, which is scene-thin by the same measure.
Our view: diversity matters for perception, depth matters for contact statistics of a specific material. Convenient for us, which is a reason to distrust it.
The strongest published critique of egocentric collection argues that the real axis is how far the captured data sits from the robot's action space, and that human video is furthest away because five fingers have nowhere to go on a two-finger gripper.
Our view: that is true for kinematics and may be false for force. A joint configuration carries morphology. A scalar newton value at a contact point does not, and both a parallel-jaw gripper and a five-fingered hand have a force command. If so, force is the one part of human demonstration sitting near zero distance. This is a hypothesis, and it is H5 above.
T3 learns tactile representations across 3M images from 13 sensors. If coarse scalar force cannot compose with that, our data risks being clean, correct and unused.
Our view: unknown. This is the question we most want a researcher to answer for us.
Newest first · failures included
Single-channel dry-electrode sensor, chosen over the better-documented alternative because it uses dry contacts on a strap. Gel electrodes give a cleaner signal and tell us nothing about the conditions a wearable will actually face.
Protocol and analysis plan written and frozen before the hardware arrives. Recording at 1000 Hz for EMG and 100 Hz for the force reference, on the same clock.
The glove design puts force sensors on the fingertips. That is the correct place to measure and the wrong place to put anything. A craftsman is paid by the piece. Cover the pad of his finger and he can no longer feel the material, his work slows, and the device comes off.
The sensing problem was never the hard part. The constraint is that the instrument has to disappear. Everything after this decision is downstream of that, including a harder sensing problem: the band infers force from muscle activity rather than measuring it at the contact.
Force sensors are not abandoned. They become the calibration reference, worn for short sessions rather than a full shift. There is no way to turn muscle signal into newtons without measuring real force alongside it.
82 companies in robot training data. Five already working on force. Three building tactile gloves, two building EMG bands, one of those funded at roughly 73 million dollars within five months of starting.
This changed the positioning rather than the plan. Sensing is not a defensible wedge. What is left is the task distribution, skilled trades where the material deforms, and access to the floors where that work still happens.
Contacted authors behind the main papers in this area, at Georgia Tech, ETH Zurich, MIT and IISc, with one question each and a request to tell us the thesis is wrong.
The specific questions: is force or shear the missing modality, does online adaptation supersede offline collection, and does the scene diversity finding invalidate a depth-first collection plan.
Entries are added when something changes, not on a schedule.
Two devices on one clock. Video says what the hands did. The band says how hard they pressed. Neither is useful without the other being time-aligned to it.
01
Capture
Head camera at 1080p30. Forearm band, 8 EMG channels at 1000 Hz. Both write locally.
02
Sync
LED on the band flashes three times at session start and end. Detected in video, mapped to frame index, drift corrected across the session.
03
Condition
Bandpass, notch, rectify, RMS envelope. Force reference converted to newtons using the per-unit calibration curve.
04
QC
Automated checks, then human review. A session either passes or is discarded with a recorded reason.
05
Release
Packaged in the format labs already load, with the datasheet and the limitations attached.
Every check is automated except the last. A failed session is kept with its failure reason rather than deleted, because the failure distribution is itself informative.
| Check | Passes if | Why it exists |
|---|---|---|
| Sync detected | Flash found at both ends, mapped lag under one frame | An unsynced session is unusable at any quality |
| Clock drift | Corrected residual under 20 ms across the session | Two clocks diverge over hours |
| Calibration age | Force reference calibrated within 7 days | Sensors drift under repeated load |
| Channel saturation | Under 1 percent of samples at rail | A clipped channel silently reports the wrong force |
| Dropped frames | Under 0.5 percent | Gaps break the alignment we just established |
| Hand visibility | Above a stated threshold, computed the same way as published datasets | Comparability |
| Band placement | Photo taken at donning, position logged | H2 says placement moves the mapping |
| Consent | Signed, versioned, in the worker's language, revocable | Non-negotiable |
| Human review | One reviewer confirms the task matches the label | Automated checks do not catch a mislabelled task |
Structure fixed · nothing collected yet
Following the datasheets-for-datasets convention. Published with the first release and versioned alongside it, including the parts that make the data less useful.
| Field | Planned content |
|---|---|
| Motivation | Testing whether contact force improves manipulation learning on deformable materials |
| Composition | Paired egocentric video, forearm EMG, derived force estimate, hand orientation, task and material labels |
| Collection | Real production work in operating workshops, not staged tasks. Workers paid above local shift rate. |
| Skill annotation | Years in trade per worker. Same task recorded across experience levels at the same bench. |
| Preprocessing | Stated filter parameters, calibration curves shipped per unit, raw counts retained alongside newtons |
| Not measured | Shear. Tactile texture. Absolute per-finger force with certainty, since the band infers rather than measures. |
| Known limitations | Two trades, two geographies, small worker count in v0. Sensor variance and hysteresis quantified and published. |
| Consent and revocation | Versioned consent, revocable after the fact, with the process for removing a worker's sessions |
| Licence | Open. First release free to use, including commercially. |
Consent is written, versioned, read aloud in the worker's language, and revocable after recording. Workers are paid roughly double the local rate for the hours they wear the equipment, and the rate does not depend on what the data turns out to be worth. A physical stop control ends recording immediately, without a phone and without asking anyone.
This work instruments people at their place of employment, where slowing them down costs them money. Treating that as a procurement detail rather than a research ethics question would be a mistake, and any lab using the data will want to know how it was obtained.
We use language models heavily and it is worth being specific about where they help and where they do not, since the failure modes matter more than the wins.
Annotated · what it claims, what it costs us
Nair · 2026
The strongest argument against what we are doing. Proposes action-space distance as the real axis and places egocentric human data furthest from it. Read this before anything supportive.
2026
Bimanual tactile estimation from egocentric video. If this holds at production accuracy on our task distribution, the band is unnecessary and our position becomes ground-truth provider.
2026
The counterweight, with numbers. Vision-only 31 percent, plus tactile 54 percent, all modalities 80 percent across five contact-rich tasks. Robot-side sensing, not human-side.
2026
The spec we build against, and the finding that challenges our site plan. Also the sentence this company rests on: video alone cannot capture the contact signals manipulation depends on.
Georgia Tech · 2024
Why the category exists. An additional hour of human hand data beats an additional hour of robot data, with large reported gains from co-training.
Georgia Tech · CoRL 2024
Human tactile demonstrations collected directly from human hands, with an audio channel alongside. Someone solved a version of our hardware problem already.
2026
Online tactile residual learning that explicitly needs no offline tactile demonstrations. Same lead author as MimicTouch, which makes it the most interesting tension in the field.
MIT · ICRA 2025
Camera-based tactile plus an acoustic channel that hears contact beginning and breaking, which is precisely where resistive force sensing is weakest. Designed for durability, which is our constraint too.
MIT · CoRL 2024
Tactile representations across 3M images and 13 sensors. Determines whether scalar force data can compose with what the field is building, or is stranded.
RSS 2024
The only real precedent for distributed collection across sites you do not control. Read for the quality problems that only surfaced after collection had finished.
Hugging Face · ICLR 2026
The format labs actually load. We build to it from the first session, because converting later costs adoption.
Kept current, including when there is nothing to report.
| Item | State |
|---|---|
| Hours collected | 0 |
| Preregistration | v0.1 published, before collection |
| Force reference | sensors in hand, calibration in progress |
| Band | single-channel pilot, hardware ordered |
| Sites | two workshops in conversation |
| Funding | self-funded |
| First release | not scheduled, will be open when it happens |