August 20, 2026
DRG#06: Anthropic Robotics Research - Visual Perception Tools for Teleoperation
Participants: anurajenp, sn37, rafa_0x
The DRG group convened to discuss Anthropic's robotics research, specifically examining how language models can be applied to robotic control and perception tasks. The meeting centered on teleoperation methodologies and the challenge of giving AI models better spatial awareness when controlling robots from an egocentric viewpoint. The team reviewed Anthropic's experimental findings on various perception assistance tools designed to help the model interpret visual information and understand directional cues in real-time robotic operations.
- Anthropic tested multiple visual aids to improve model perception in robotics, including crosshairs, depth heatmaps, third-person cameras, and orientation compasses.
- A compass tool providing orientation in degrees significantly outperformed other visual assistance methods for helping the model understand directional context.
- The research focuses on enhancing Claude's ability to understand spatial relationships and navigate environments through teleoperation.
Session Recording Summary · 46m 16s · Full notes ↗
Reading: The primary reading was Anthropic's report/blog post on using LLMs to control robots (referred to as "the Anthropic one" / "the first link"), alongside a second linked piece described as "the drone link." Exact titles were not stated in the transcript. The discussion centered on Anthropic's experiments testing LLMs at various levels of robot control (limb control, task planning, navigation) across different robotic platforms and models.
The group went around sharing reactions to Anthropic's robotics/LLM experiments, then broadened into discussion of the practical stack for LLM-controlled robots — teleoperation, latency, and where compute should live. Jenna introduced an organizational note about a planned PI (Protocol Institute) quarterly journal and possible written/interactive outputs from the group. The session ended with Anuraj demonstrating a live browser-based teleoperation setup controlling his robot in Finland.
- **rafa:** Robotic control by LLMs feels increasingly like a "solved problem" in principle — the open questions are *when* and *which narrow use cases* get solved first. Noted a recent general AI model for robotics that learns by watching examples. Raised the "stack of problems" that remain even if driving/control is solved: latency, real-time reconstruction, and deciding what compute lives on the robot vs. off it. Anchored this in his own struggle to get reliable home Wi-Fi/mesh networking.
- **shreeram:** Wants to see the actual implementation code, not just the writeup (his robotics background is minimal). Found it notable that in the "policy task," supervision made older models *worse*. Confused by unexplained ~20-second gaps in the pendulum/classic-control tasks. His main critique: the paper never establishes what the *best non-LLM method* can achieve as a baseline, so it's unclear whether LLMs are actually the right tool or just a hammer being applied everywhere.
- **Spencer:** Most interested in the experiment-level details — how small setup changes yield outsized improvements. Highlighted two examples: adding a "compass" giving the locomotive robot orientation in degrees, and adding a crosshairs/reticle to the camera view. Argued we should look for ways to help models "in the way they want to be helped," rather than forcing them into human-style control paradigms.
- **Anuraj:** Had a "meh" reaction. Argued the path of using a language model alone won't yield much for real-world navigation — you need a *world model* for embodied action (humans don't verbally reason about walking or holding a cup; it's a different mode of thinking). Thinks LLMs are better suited to higher-level planning, while low-level control needs models trained from scratch with embodiment. Agreed with shreeram's "hammer" framing — Anthropic seems to try many reward setups to see what sticks. Suspects robotics isn't a major Anthropic priority, hence the simple experiments.
- **Maier:** Read the goal as a mix of *safety* (will people use our AI to build killer robots?) and *capability assessment* (how good are LLMs today?). Noted they tested LLMs at ~four levels of abstraction but tried to solve all layers with one model — whereas he'd use a *different LLM per layer*, both because that mirrors biology and because these problems are too hard to solve simultaneously in one decision. Called Anuraj's "let LLMs do what they're good at, glue from elsewhere" view exactly the kind of prescription the paper lacked. Found "points of light": LLMs writing their own control code, and the question of what *protocol* a robot should use to get help from an LLM (given LLMs are energy/CPU heavy — maybe a robot runs a small/pretrained model locally).
Questions & Disagreements: - **Is LLM-only control the right approach?** Anuraj (and to a degree shreeram) are skeptical for low-level, real-world navigation, favoring embodied/world models; Spencer and rafa are more optimistic about incremental improvements and helping the model in its own terms. Not resolved. - **Missing baseline:** shreeram's open question — what's the best a non-LLM method achieves, and is LLM control e
Participants: Anuraj R, rafa (UTC+1), shreeram, Spencer Norwick U-0700, Maier, 🙊 Jenna Dixon, Giovanni Merlino