Skip to content
  1. Home
  2. AI & Machines
AI & Machines

On-device inference gets good enough to change product decisions

When a model fits in 3 watts, the interesting question stops being accuracy and becomes what you no longer have to send anywhere.

On-device inference gets good enough to change product decisions
On-device inference gets good enough to change product decisions — XENVORA Photograph: Computer processor close up, Rawpixel (CC0) · Licence

Small models running locally used to be a compromise you apologised for. This generation of accelerators removed the apology for a broad class of tasks: transcription, translation, summarisation and image cleanup.

The privacy dividend is accidental

Teams did not move inference to the edge for ethics. They did it for latency and cloud bills. Not shipping user data anywhere turned out to be a free side effect worth marketing.

The tradeoff is update cadence. A cloud model improves the week the weights change; a device model improves when users install a firmware update, which some never do.

Our advice to product teams: put the fast, private path on device and keep a cloud escape hatch for the 5% of requests that genuinely need a larger model.

Omar Adel
Machine Learning Correspondent
Omar Adel

Omar reports on applied machine learning, with a stubborn habit of asking vendors for reproducible benchmarks.

Reader comments 0 comments

Sign in to comment.

No comments yet. Be the first.