AI and LLMs · Runtime

ONNX Runtime

Microsoft's open-source engine for running trained AI models saved in the ONNX format, on servers, on phones, in Node.js and inside a web browser.

Run it yourself · Open source · updated

How it works

ONNX (Open Neural Network Exchange) is an open file format for trained models, so a model built in PyTorch or TensorFlow can be exported once and run elsewhere. ONNX Runtime loads that file and runs it efficiently on whatever hardware is present, through plug-in execution providers for CPUs, NVIDIA GPUs (CUDA, TensorRT), Windows GPUs (DirectML), Apple devices (Core ML) and more.

It suits smaller, focused models rather than big chat models: text embeddings for search, classifiers, image recognition, OCR and speech. onnxruntime-node runs them inside a Node.js server, and onnxruntime-web runs them in the browser with WebAssembly or WebGPU, so data can stay on the user's device. Hugging Face's Transformers.js is built on it and handles downloading and preparing models for you.

ONNX Runtime pricing

Open source

Free (MIT).

Approximate, checked September 2026.What the other tools cost

More in AI and LLMs

Run it yourself

All 22 AI and LLMs terms

Crafted in the dark. Shipped to the world.

Tell us what you are building. You get a private project space with a proposal and a line-by-line quote within a day.