Small models running locally used to be a compromise you apologised for. This generation of accelerators removed the apology for a broad class of tasks: transcription, translation, summarisation and image cleanup.
The privacy dividend is accidental
Teams did not move inference to the edge for ethics. They did it for latency and cloud bills. Not shipping user data anywhere turned out to be a free side effect worth marketing.
The tradeoff is update cadence. A cloud model improves the week the weights change; a device model improves when users install a firmware update, which some never do.
Our advice to product teams: put the fast, private path on device and keep a cloud escape hatch for the 5% of requests that genuinely need a larger model.
Reader comments 0 comments
Sign in to comment.
No comments yet. Be the first.