The AI agent for inference optimization that actually understands your target hardware

Squeeze more model performance out of embedded compute (like NVIDIA Jetson) than you realize is possible. Upgrade your onboard perception/autonomy, without hiring more optimization specialists.

Built for ML teams in

Autonomous VehiclesRoboticsSmart CamerasDrones
RunLocal EnvironmentSelf-hosting possible toeliminate IP security concernsRunLocal Web UIFor you to auditthe agent's workYour Trained Model RepoCode · Weights · Validation DataOnly what's required for inferenceoptimization and testingOff-the-shelf coding agentCodex · Claude · GeminiRunLocal CLIFor the agent touse our engineRunLocal EngineHardware Performance ModelingOn-Device OrchestrationExperiment LedgerYour Real Target HardwareJetson Orin/Thor · Qualcomm

The RunLocal Engine

Proprietary performance modeling on your target hardware and on-device orchestration infra (queuing, environment config, etc.), which enable your coding agent to truly understand what's possible with your AI stack on your target hardware, and unlock performance that you wouldn't get otherwise.

Secret Sauce In The RunLocal Engine

Statistical performance modeling grounded in benchmark data we've already collected across real devices, on top of execution infra nobody wants to build. Generic coding agents will never have this, and the value grows as they become generally smarter.

Hardware Performance Modeling

Custom analysis of your model, based on real benchmarks for your target hardware, which feeds into a causal understanding of how model changes might affect on-device performance - clarifying your model's bottlenecks, optimization opportunities and performance ceiling.

On-Device Orchestration

Infrastructure that ensures an agent can safely organize, execute and observe work asynchronously on your target hardware (queuing, HW/SW environment config, etc.)

Experiment Ledger

Every measurement is bound to what produced it (code change, software versions, input data, etc.) in an append-only record, published by the system not the agent. So you can audit any claim, reproduce any result, and every experiment compounds into knowledge the agent can learn and iterate on.

How You Use It

Install and run our CLI to launch your coding agent in the RunLocal environment, connect it to your real models and target hardware, and prompt it like you normally would — our infrastructure does its magic under the hood.

1

Integrate

Connect RunLocal to your model repos (code, weights and validation data) and your real target hardware.

Self-hosting our software in your infra is possible to eliminate IP/security concerns.

2

Optimize

Run our CLI to launch your coding agent in the RunLocal environment, then prompt it exactly like you normally would.

Bring your own AI vendor and API keys (e.g. Codex or Claude).

3

Audit

Use our Web UI to track the agent's experimentation and verify its output, in real time as it works.

Value To Your Business

More optimized models, shipped sooner, on cheaper compute onboard, with a leaner team.

Better

More Optimized Models

Lower latency and memory (same accuracy), or better models with less compute.

Faster

Ship Faster

Hit performance targets in days instead of weeks, or hours instead of days.

Cheaper

Less Headcount & Compute

Avoid hiring rare optimization experts. Downgrade your onboard compute.

Coding Agents Alone Aren't Enough

These failure modes don't go away as coding agents get more capable at generic coding — closing them takes specialized infrastructure built for on-device optimization.

Unreliable On-Device Benchmarking

Benchmarking jobs collide, jobs silently crash, and corrupted runs quietly hinder experimentation.

What it takes

Trustworthy on-device benchmarking, which seamlessly deals with many parallel agents, requires a sophisticated benchmarking system.

Shallow Optimization Hypotheses

Basic optimization hypotheses because they don't have a deep understanding of your target hardware.

What it takes

Cause-and-effect must be derived from benchmarking on real hardware with a specific causal analysis system; better generic reasoning isn't a substitute.

Cheating & Non-Trivial Verification

They find ways to hit performance targets by breaking real constraints, and it's non-trivial to verify results.

What it takes

Changes/results must be precisely recorded/validated independently, and a purpose-built UI is needed for inspecting results and artifacts.

Reduce Costly Bottlenecks

Even with today's coding agents, model inference optimization still drags you into the same grind. With RunLocal, the agent actually handles them for you.

Performance Bugs

Poorly supported layers, unexpected issues after quantizing, and other silent-but-deadly surprises you still end up chasing down alongside your agent.

Endless Trial-and-Error

Babysitting the agent through attempt after attempt, re-explaining context, and hand-holding it toward something that actually hits your numbers.

Missed Performance Gains

Not knowing whether you're near the hardware's limit or leaving speed on the table — and no way to tell if another round of optimization is worth it.

A Continuous Optimization Loop

Your models and validation data go in. The agent hypothesizes, transforms, compiles and benchmarks on real hardware — learning each round until it hits your on-device performance targets.

PyTorch
ONNX
+ Validation Data & App Code
(e.g. Pre/Post-Processing)
Hypothesize
Transform
Compile
Benchmark
Learn
RunLocal
NVIDIA
Jetson OrinJetson Thor
Qualcomm
+ Ambarella & TI soon

Backed By

468 Capital
Y Combinator
Ritual Capital

and more

Frequently Asked Questions

Things you might want to know before trying RunLocal