Apache 2.0 · self-hosted · ROS 2 and custom edge stacks

Your logs said the
robot was healthy.
It stopped for a shadow.

Watchpoint is AI failure forensics for physical AI. It captures what your model saw, what it predicted, and what your policy decided at the moment of failure — then replays that exact inference. Root cause at the AI layer, not just the logs.

No signup for the demo. Self-hosted install runs entirely on your own infrastructure.

Four things your stack throws away every frame — and needs at 2am:

01
What the model saw
Synced camera, lidar, and depth frames at inference time
02
What it predicted
Outputs, per-class confidence, attention and saliency
03
What the policy decided
Chosen action, ranked alternatives, and their scores
04
Whether the input was novel
Embedding distance from your training distribution

The metrics are green.
That's the whole problem.

An AMR halts mid-aisle. You open the dashboards. CPU is at 40%. Memory is flat. Thermals are nominal. Every ROS 2 node is publishing at its nominal rate. Nothing threw an exception.

So you pull the rosbag, scrub through it by hand, and eventually find the frame — a hard shadow across a loading bay that the detector called an obstacle at 0.71 confidence. Three days gone.

Infrastructure monitoring is structurally blind to this. It is built on the assumption that resource health predicts failure. At the AI layer that assumption is exactly inverted: the machine is perfectly healthy, and the model is wrong.

Datadog, Grafana, Prometheus
Was the machine healthy?
Green. Unhelpfully.
Foxglove, rosbag tooling
What did the sensors publish?
Everything, if you know what to look for.
Sentry and APM
Did the code throw?
No. It ran perfectly and returned a wrong answer.
Watchpoint
Was the model right — and if not, why?
AI-002: input 2.7σ out of distribution.

Eight ways an AI system fails silently

Watchpoint names the failure instead of handing you eleven charts. Each rule runs against captured model state, not just system telemetry.

AI-001
Perception confidence collapse
shippedhigh

Detection confidence p50 over 60s drops more than 30% from baseline

AI-002
Out-of-distribution input
shippedmedium

Embedding distance exceeds 3σ from the training-set centroid

AI-003
Inference latency spike
shippedmedium

p99 latency over 60s exceeds 2x baseline

AI-004
Per-layer latency anomaly
low

A single layer exceeds 5x its baseline latency

AI-005
Decision-perception mismatch
high

Policy chose an action incompatible with a high-confidence detection

AI-006
Attention drift
low

Attention center-of-mass shifted more than 50% of frame from baseline

AI-007
Output saturation
medium

Softmax entropy below 0.1 nats across diverse inputs

AI-008
Sensor degradation upstream of model
medium

Image sharpness or lidar density dropped more than 40% from baseline

Rules marked shipped run today. The rest are specified and on the roadmap — we don't claim what isn't merged.

How it works

Instrument once, capture continuously in memory, keep only what matters.

Step 1

Instrument

Two lines attach forward hooks to your PyTorch model. The collector rings a fixed-size buffer in-process — designed for under 1% overhead at p99, with nothing transmitted until an incident fires.

  • PyTorch adapter (shipped)
  • ONNX Runtime and TensorRT (roadmap)
  • ROS 2 topic, node, and lag monitoring
  • Go edge agent for host metrics
Step 2

Correlate

On an incident trigger, the buffer flushes and joins the model timeline to system telemetry, ROS 2 topic health, and the deployment that was running — matched by weights hash.

  • Model, sensor, and host state on one timeline
  • 7 system rules plus the AI rule engine
  • Incidents grouped by release and weights hash
  • Optional two-sentence LLM summary
Step 3

Replay

Export a portable bundle any engineer can open, or re-run the captured inputs against new weights to prove the fix before it reaches the fleet.

  • Replay bundle ZIP export (shipped)
  • Deterministic replay sandbox (roadmap)
  • Attention overlay on the failure frame (roadmap)
  • Decision trace with ranked alternatives
Self-hosted by default

Your camera data never leaves your VPC.

Footage from a customer's warehouse is usually contractually un-exportable. That single fact kills most observability vendors in robotics procurement, so we built for it from the start.

Watchpoint runs entirely on your infrastructure. Model weights are hashed for lineage, never uploaded. LLM summaries are optional — with no API key configured, the rules engine degrades to deterministic text and every feature keeps working.

Runs in your infrastructure
Docker Compose today; your own Postgres, your own storage.
No outbound dependency
No telemetry home. The stack functions fully air-gapped.
Weights hashed, not uploaded
Incidents tie to a weights hash so you can group by release.
Apache 2.0 core
Every collector is open source. Read it before you put it on a robot.

Run the whole stack locally

Clone, compose up, seed. Three demo incidents, each carrying both system telemetry and captured AI-layer inferences — no account, no cloud.

$ git clone https://github.com/sagarbpatel31/watchpoint.git
$ cd watchpoint/deploy/docker-compose && docker compose up -d
$ curl -X POST localhost:8000/api/v1/seed/demo