Watchpoint is AI failure forensics for physical AI. It captures what your model saw, what it predicted, and what your policy decided at the moment of failure — then replays that exact inference. Root cause at the AI layer, not just the logs.
No signup for the demo. Self-hosted install runs entirely on your own infrastructure.
Four things your stack throws away every frame — and needs at 2am:
An AMR halts mid-aisle. You open the dashboards. CPU is at 40%. Memory is flat. Thermals are nominal. Every ROS 2 node is publishing at its nominal rate. Nothing threw an exception.
So you pull the rosbag, scrub through it by hand, and eventually find the frame — a hard shadow across a loading bay that the detector called an obstacle at 0.71 confidence. Three days gone.
Infrastructure monitoring is structurally blind to this. It is built on the assumption that resource health predicts failure. At the AI layer that assumption is exactly inverted: the machine is perfectly healthy, and the model is wrong.
Watchpoint names the failure instead of handing you eleven charts. Each rule runs against captured model state, not just system telemetry.
Rules marked shipped run today. The rest are specified and on the roadmap — we don't claim what isn't merged.
Instrument once, capture continuously in memory, keep only what matters.
Two lines attach forward hooks to your PyTorch model. The collector rings a fixed-size buffer in-process — designed for under 1% overhead at p99, with nothing transmitted until an incident fires.
On an incident trigger, the buffer flushes and joins the model timeline to system telemetry, ROS 2 topic health, and the deployment that was running — matched by weights hash.
Export a portable bundle any engineer can open, or re-run the captured inputs against new weights to prove the fix before it reaches the fleet.
Footage from a customer's warehouse is usually contractually un-exportable. That single fact kills most observability vendors in robotics procurement, so we built for it from the start.
Watchpoint runs entirely on your infrastructure. Model weights are hashed for lineage, never uploaded. LLM summaries are optional — with no API key configured, the rules engine degrades to deterministic text and every feature keeps working.
Clone, compose up, seed. Three demo incidents, each carrying both system telemetry and captured AI-layer inferences — no account, no cloud.