AI & Analytics
A frontier-class model on a 128GB machine
Ox Alpha topped OpenRouter before anyone knew who built it. It turned out to be an open-weight model you can run inside your own perimeter, which matters more to a regulated enterprise than any benchmark.
Every so often a model shows up on OpenRouter with no name attached to it, and the guessing starts before the benchmarks do. That is what happened this month with a listing called Ox Alpha.1 Within days it climbed to the top of OpenRouter's usage rankings, the kind of adoption that normally takes a known lab and a marketing budget. Nobody knew who had built it.
The reveal came a few days later. Z.ai confirmed to Bloomberg that Ox Alpha was GLM-5.3-Flash, the newest release in its GLM line, and open-sourced it under the MIT license.2 GLM has been one of the stronger open-source families for a while, so the lineage checks out. Even before anyone knew whose model it was, the specs stood out. A context window near a million tokens. Text, image, and video input in one model.4 And because it shipped open-weight, it did not stay behind an API. Within days, people had it quantized to a size that runs on a machine with 128GB of RAM.3
Where it sits on the local AI spectrum
We spend a fair amount of time at Greyquill testing where local AI can and cannot carry its weight. On one end of that spectrum is Needle 2 from Cactus Compute, built for extremely lightweight, on-device tool calling. Small and fast, narrow by design. On the other end sits Gemma, Google's family of open models, including the variants built to run on-device. Both are useful. Both trade away capability to get there.
Ox Alpha does not fit neatly on that spectrum. A well-specced workstation can now run something close to frontier-model reasoning, rather than a lightweight assistant that happens to work offline. The gap between local and capable, and local and actually competitive, has been narrowing for a while. This is the clearest sign yet of how fast it is closing. The 128GB number is the one we keep returning to. Two years ago that bought you a small model with big compromises.
Why a regulated enterprise should care
Token costs are real money at scale, and every team running agents in production feels it eventually. Run the model on hardware you already own and that bill goes away, along with the rate limits and the dependency on someone else's uptime. That alone makes local models worth a serious look.
For a bank, an insurer, or a telco, there is a bigger reason. Data that never leaves the building is a much shorter conversation with the risk office than data that transits a third party's API. A capable model you can run inside your own perimeter turns a governance headache into a deployment decision. That does not make the auditability questions go away, but it changes who you have to answer them with.
Data that never leaves the building is a much shorter conversation with the risk office.
The harness matters as much as the model
The harnesses around these models are changing even faster than the models. Tool use is more reliable than it was a year ago, and the repetitive steps that used to need someone watching over them now mostly run unattended. Pair a capable open-weight model with a solid harness and it starts closing the gap with a hosted frontier model on real work.
We think that pairing is where a lot of the next real progress in applied AI comes from. Local models and the harnesses that make them useful, compounding month over month. 2027 is shaping up to be the year local AI and agent harnesses stop being two separate conversations. That is our bet.
One caution from our own testing. A model on your own hardware does nothing about the data underneath it. If the records it reads are inconsistent, unlabelled, or unowned, you have moved the problem indoors and made it cheaper to hit. The harness is only as good as what it can trust, and that is a data question before it is a model question.
What to do with this
You do not need to move a workload local tomorrow. You do need to know what one of these models can do on your own data before your next infrastructure decision assumes it cannot. A weekend on a 128GB machine will tell you more than a quarter of vendor briefings.
Curious what is holding other teams back from testing one on real work: the hardware, the governance questions, or the state of the data it would have to read.
