Calibration and evaluation method

Reviewed · Zagolo

A Gap Audit is Zagolo’s scoped robot dynamics calibration pilot. We fit selected model parameters from hardware logs, then compare the original and fitted models on complete runs reserved before fitting. The comparison evaluates motion prediction, with improvements, regressions and unresolved errors reported together.

We are seeking our first pilot partners. This page explains the proposed evaluation method; it is not a published hardware benchmark or a claim of demonstrated improvement. The robot, model, controller and available recordings determine what can be identified.

Start with a repeatable mismatch.

Useful starting points include a joint that lags, a reversal that sticks, or a trajectory that drifts when commands are replayed in simulation. We review the setup and recordings before agreeing which parameters to investigate. Friction, damping and actuator delay are examples, not a guarantee that every parameter is identifiable in every dataset.

We record the robot configuration, payload, firmware, controller and simulator/model versions. Commands and measured responses must have clear units, joint mappings and timing. The robot recording guide describes the signals and run notes to preserve.

Agree the comparison before fitting.

We establish the original-model baseline and agree how motion prediction will be scored. Both models receive the same replayed commands, starting state and relevant simulator settings. Stream alignment, any interpolation, and treatment of missing or invalid samples are documented.

The protocol identifies which runs are used for fitting and which complete runs are reserved for final evaluation. It also records the parameters, bounds and objective used for fitting. The test set is not used to select parameters or tune the model.

Fit only the agreed parameters.

The fitting work changes the selected model parameters using the calibration runs. We keep a record of original and fitted values, units and configuration changes. A good fit on calibration data is a diagnostic result; the reserved-run comparison is the evaluation that follows.

Some parameters can compensate for each other, especially when recordings contain limited excitation. A closer trajectory match alone does not establish that the physical parameters are uniquely identified. The report explains limitations in the data and remaining plausible sources of mismatch.

Freeze the model and evaluate once.

After fitting, we freeze the model and compare both versions on the same reserved runs. We report every final test result, including runs or joints that regress. If evaluation reveals a problem that motivates further fitting, that becomes a new experiment with a new untouched test set; it is not silently substituted for the original result.

Configuration changes, contacts, actuator saturation or timing faults can make a comparison misleading. We preserve run notes and state any exclusions and their reasons. We do not select only the best trajectory segments.

Read error measures in their original units.

For aligned joint-position samples, let eᵢ = q̂ᵢ − qᵢ, where q̂ᵢ is the model prediction and qᵢ is the hardware measurement. Three useful diagnostics are root mean square error RMSE = √(Σeᵢ² / N), mean absolute error MAE = Σ|eᵢ| / N, and maximum absolute error max|eᵢ|.

These retain the joint’s position units: radians for a rotational joint, or meters for a translational joint when those are the recorded units. We report results per joint and per run. Mixed-unit errors should not be averaged into an unexplained overall score.

A percentage reduction can be useful alongside the actual errors, but is undefined when the baseline error is zero. Sample counts, timing alignment, and the same set of observations must accompany the comparison. The agreed pilot may require additional measures for its particular mismatch.

What a useful evidence record contains

Motion prediction and policy transfer

A closer dynamics fit may help a team decide what to investigate before another policy-training cycle. It does not establish task success. Perception, contact behavior and task design can still limit transfer. A policy test on hardware is a separate experiment.

The deliverables are an updated model for the agreed workflow, a before-and-after comparison and a review of the remaining mismatch with the engineer doing the work. Scope, robot time, delivery schedule and fixed fee are agreed in writing after setup review.