Tasks

This section provides a overview of the tasks evaluated in the paper.

point-to-disk

A velocity-controlled point-mass with state \(\rvx\) (cartesian position) must be driven into a target disk of radius \(r^* = 0.01\) m at the origin and held there. Given a velocity command \(\rvu\), the system evolves as \(\dot{\rvx} = \rvv\) with \(\rvv \sim B_\epsilon(\rvu) := \{ \rvv : ||\rvv - \rvu||_2 \leq \eps ||\rvu||_2 \}\) (\(\eps = 0.25\) in our experiments). The actuation noise grows with the commanded speed, so slowing down (\(\rvu \to \mathbf{0}\)) allows the controller to suppress actuation noise. The controller, operating at 50 Hz, is scored by its dwell time inside the target for rollouts of length \(H = 500\) steps (10 s) from states on the unit circle: \[\rmJ(\pi) = \E_{\pi,\,\{\rvx_0 : ||\rvx_0||_2 = 1\}}\left[ \sum_{t=0}^H \Ind{||\rvx_t||_2 < r^*} \right]\]

Illustration of point-to-disk dynamics

Fixing a coordinate system XY, the robot can move at most 1 m/s in either direction, so the action space is \(\gA = [-1, 1]^2\). Figure 1 illustrates a quiver plot of an oracle policy that has access to the perfect state \(\rvx\). The policy commands the maximum possible velocity towards the target when outside the disk, and starts slowing down to zero as it enters the target disk. The oracle policy achieves a performance of \(\rmJ_\text{oracle} = 456.23 \pm 4.59\) (\(N=1000\)).

Figure 1: Quiver plot of oracle policy on point-to-disk task.

Sensors The designer can equip the various controller modes with the following sensor types. Each sensor reveals a different slice of the robot’s position \(\rvx\), as illustrated in Figure 2.

  • Cartesian (GPS): noisy position \(\rvx + \gN(0, \sigma_\text{xy}^2 \mI_2)\), with \(\sigma_\text{xy} \in \{0,\,10^{-2},\,5{\cdot}10^{-2},\,10^{-1},\,5{\cdot}10^{-1}\}\).

  • Radial & Sector (RadialSector): a discretized polar reading returning the active sector among \(s \in [1,360]\) equiangular sectors (optionally offset by \(\phi\) radians) and the active radial band among \(\text{len}(\rvd){+}1\) bands defined by increasing distance thresholds \(\rvd = (d_1, \dots) \le 1\).

    • The per-step output is a concatenation of the one-hot over radial bands (empty when \(\rvd = ()\)) and unit-vector \((\cos\theta_s,\sin\theta_s)\) corresponding to centroid of the active sector.
    • There is some Gaussian noise in measuring the polar coordinates of the robot \((r, \theta)\).
      • The noise scales are tunable with \(\sigma_\theta \in \{0,\,10^{-1},\,2{\cdot}10^{-1},\,4{\cdot}10^{-1},\,8{\cdot}10^{-1}\}\) and \(\sigma_r \in \{0,\,10^{-2},\,5{\cdot}10^{-2},\,10^{-1}\}\).

Mode Transition Observation The mode transitions are based on the robot’s position vector \(\rvx\), which corresponds to the Markov state in this task.

The designer’s goal is to meet a performance target \(\rmJ_\text{target}\) while minimizing the sensing cost incurred.

Figure 2: Interactive visualization of sensors for the point-to-disk task. Configure the parameters of RadialSector (left) and GPS (right) and click inside the unit circle (dashed line) to probe the robot’s instantaneous observation. The visualization depicts the costs of the configurations under the cost-structures CheapGPS and CostlyGPS on the top. The cost difference between the RadialSector (left) and GPS (right) configurations is depicted in the center, with a RED background indicating a cheaper RadialSector configuration.

RF-DMC

We construct rangefinder variants of three dm_control tasks — cartpole-swingup, cup-catch, and finger-spin. In each, the agent observes the world through a RayScanSensor consisting of 360 equiangular rangefinders, spanning \(360^\circ\), rigidly attached to a task-relevant body site (pole tip, cup center, fingertip). This rangefinders report the distance to the nearest surface as visualized in Figure 3. The rollouts and performance of the oracle policies on these tasks are visualized in Figure 4.

(a) cartpole-swingup
(b) cup-catch
(c) finger-spin
Figure 3: Visualization of the RF-DMC tasks and the RayScanSensor observations.

cartpole-swingup

Rollout:

cup-catch

Rollout:

finger-spin

Rollout:

\(\rmJ_\text{oracle} = 881.8 \pm 4.9\)

\(\rmJ_\text{oracle} = 977.8 \pm 15.7\)

\(\rmJ_\text{oracle} = 969.9 \pm 11.7\)

Figure 4: Rollouts of oracle policies on the RF-DMC tasks. The performance is reported over 1000 rollouts.

Rangefinder noise model We adopt the popular mixture-of-noise model (Probabilistic Robotics, Thrun et al.) to corrupt true range measurements \(z^*\), in the range \([z_\text{min}, z_\text{max}] = [0, 10]\) m, with the following noise components:

  • hit — measurement noise, \(\gN(z^*, \sigma_\text{hit}^2)\), truncated gaussian located at the true range, clipped to \([z_\text{min}, z_\text{max}]\) (weight \(c_\text{hit}\))
  • short — interfering obstacles, \(\propto \text{Exp}(\lambda_\text{short})\) on \([0, z^*]\) (weight \(c_\text{short}\), held at \(0\) across grades in our experiments)
  • max — missed detections, \(\delta\)-spike at \(z_\text{max}\) (weight \(c_\text{max}\))
  • rand — random noise, uniform over the measured range \([z_\text{min}, z_\text{max}]\) (weight \(c_\text{rand}\)).

The designer tunes discrete sensor quality levels \(\in \{\texttt{A}, \dots, \texttt{F}\}\) that set operating specific mixture weights and measurement noise scales \(\sigma_\text{hit}\). The per-step sensor cost is \(\big(\text{energy}(\texttt{quality}) + \tfrac12 \log_2 H\big)\cdot n_\text{rays}\), where the per-ray grade energy is obtained from the SNR for the chosen rangefinder noise model, \(H\) is the observation history length, and \(360 / n_\text{rays}\) is the angular resolution. Figure 5 presents an interactive visualization of the noise model and sensor costs for the RF-DMC sensor configurations.


Figure 5: Interactive visualization of the Rangefinder noise model. Left: For the selected sensor quality (A–F) and true range \(z^*\); the panel below reports the grade’s mixture weights and measurement statistics. Center: visualizes the density \(p(z \mid z^*)\). Right: reports the per-step sensor cost of the configuration.

clutter-nav

A disk-shaped robot of radius \(0.5\) m with unicycle (differential-drive) dynamics must reach a goal while avoiding obstacles in a cluttered \([-15, 15]^2\) m plane. At each reset, \(n_\text{static} \in [5, 20]\) static and \(n_\text{dynamic} \in [0, 15]\) dynamic disc obstacles (radii \([0.3, 2.0]\) m) are spawned at random, and a goal is sampled \(10\)–\(15\) m away. The ego state is \(\rvx = (x, y, \theta)\) and the action \(\rvu = (v, \omega)\) commands forward velocity \(v \in [0, 5]\) m/s and turn rate \(\omega \in [-3, 3]\) rad/s; dynamics integrate at \(dt = 0.05\) s (20 Hz) with command-scaled actuation noise. Dynamic obstacles follow their own goal-seeking + repulsion policy toward periodically resampled waypoints. Each mode jointly configures the Radar sensor and an MPPI planner that acts under certainty equivalence — treating the noisy detections as ground truth state.

Radar sensor The agent perceives the world through a radar-like sensor that reports detected obstacles in the ego-centric frame, each detection is a tuple \((\evr, \evtheta, \evs, \evv)\) — range, bearing (relative to heading), obstacle size, and radial velocity. Only obstacles within the specified range (under max_range) and field of view (\(|\evtheta| \le \texttt{fov}/2\)) are registered. These detections are subject to noise as outlined below:

  • Detection probability The probability of detecting the obstacle drops with the distance via the SNR proxy \(P(\evr) = (R_0 / \evr)^4\), as \(p_\text{detect} = p_0^{\,1/(1+P)}\), which provides reliable detection for \(r \ll R_0\) but decays to \(p_0 = 10^{-5}\) as the range increases.
  • Occlusion: Detections respect visual occlusions. The ego-centeric detections are quantized into \(1^\circ\) bins and the nearest obstacle in each bin is the only obstacle registered.
  • Range-scaled Noise Noise scales with the distance across channels (range, size, radial velocity) as \(\sigma_*(\evr) = \hat\sigma_*\, \evr^2 / R_0\), with fixed base scales \(\hat\sigma_r = 0.2\) m, \(\hat\sigma_\text{az} = 0.05\) rad, \(\hat\sigma_\text{sz} = 0.08\) m, \(\hat\sigma_v = 0.3\) m/s.
  • Spurious Detections: Spurious detections are registered at \(\text{Poisson}(1)\) per scan, these are uniform over the sensor’s coverage.

Of the various knobs, the designer controls only \((R_0, \texttt{max\_range}, \texttt{fov})\) parameters of the sensor — larger values give cleaner, longer-range, wider coverage at higher cost. Figure 6 presents an interactive visualization of the Radar sensor’s noise characteristics. Figure 7 visualizes rollouts of a maxed-out sensor and planner configuration MPPI controller.

Figure 6: Interactive Radar sensor visualization for the clutter-nav task. Left: a randomly sampled world scene — the ego robot (teal, tick = heading) with its field of view and range disk, static obstacles (gray), dynamic obstacles (terracotta, with velocity arrows), and the goal (navy dashed). Right: the instantaneous radar observation in the ego-centric frame, where each detection is colored by the radial velocity and scaled by measured obstacle size; dashed markers are false positives. Click Sample scene to sample a new configuration, or Sample observation to resample a (stochastic) radar scan on the current scene. The control panel exposes the three designer-controlled knobs \((R_0, \texttt{max\_range}, \texttt{fov})\) and reports the resulting per-step sensor cost — the energy term \(w_s(R_0/R_\text{ref})^2\) plus the coverage term \(w_c\,\tfrac12\,\texttt{max\_range}^2\,\texttt{fov}/A_\text{ref}\) — using the weights \((w_s, w_c) = (1.0, 0.5)\).

Figure 7: Rollouts of MPPI controller with maxed-out sensor and planner configurations. Left: the world scene view Right: corresponding ego-frame detections.