← Writings

Building Glider: Lessons from a Computer Use Agent

Why I Built Glider

I found a gap in AI capabilities; that's why. AI has reportedly solved one of the million-dollar Millennium Prize Problems. Yet if you gave it Pong, the 1970s video game, it wouldn't be able to keep up with you. This makes sense considering large language models are the main format of AI today. In fact, “giving the AI a game to play” isn't even really a well-defined statement. LLMs respond to text with text. But over the past year or two, developers started prompting these large language models to control the mouse and keyboard: the computer use agent was born.

I started working on computer use agents during my final semester at Carnegie Mellon University while studying multimodal machine learning, and the standard architecture struck me as vastly inefficient. These agents work by repeatedly sending screenshots to massive LLMs asking for executable code. This summer, I took it upon myself to build a blazingly fast computer use agent, one that would not only be able to play a game like Pong, but also work on everyday tasks at incredible speeds. I quickly learned about the limitations though…

What Is a Computer Use Agent?

I want to be absolutely clear on what a computer use agent is, since the definition isn't obvious, even to my friends who study computer science. A computer use agent operates your screen, mouse, and keyboard. Think of remote desktop software, where another person controls your mouse and keyboard from their own computer. A computer use agent works the same way, except there's AI on the other side, and you're the one telling it what to do.

Don't confuse them with other types of “AI agents” like OpenClaw or Meta Muse, because those work very differently. They get work done in the background using commands, APIs, and remote virtual machines. You text them what to do and they autonomously get it done. That's not literally operating your keyboard and mouse. The flexibility of computer use agents means they work in environments where these background AI agents don't, because screen navigation is software agnostic.

What I Set Out to Build

I set out to build the world's fastest computer use agent. It would have a polished user interface. It would be cross-platform across Windows, macOS, and Linux. It would run at the desktop level, not the browser level. It would cost nearly nothing to run by leveraging smaller models. All of this together really was a reach goal.

By the end of the summer, I had a working prototype, fully built in Rust. It included several general improvements: structured tool calls instead of raw code, error correction, prompt engineering, a feedback system, and action caching. Ultimately though, because I was working with LLMs, I was stuck in the same looping paradigm that I was looking to break out of: send a screenshot, get a command back, run it, repeat. I just optimized it to be as efficient as possible. Was it fast enough?

The first thing I had Glider attempt was Wordle. It first-tried it, faster than I had ever seen from another agent. But it was still slow. My vision of having it complete so fast you couldn't even comprehend what was going on was far off. It wasn't the revolutionary tool I had imagined.

The Latency Wall

Over the summer, Alibaba released Qwen 3.8 27B, a small language model, and arguably the best one to fit on consumer hardware. With 3-bit quantization, I was able to get it running on my RTX 5070 Ti completely offline, allegedly giving me the power of Claude Opus 4.6 for free. With a score of 84.3% on the OSWorld benchmark, Qwen 3.8 27B was clearly built for computer use tasks. I selected it as the backend brain of Glider.

While it excelled at accuracy and reasoning, it was still so slow. To be precise, I could not consistently get the prefill token speed faster than 1500 tok/s or the output faster than 100 tok/s. Considering a screenshot is at least 2000 tokens and the output is about 100 tokens, this meant at least 2 seconds each iteration. Datacenter GPUs might have been fast enough but would have cost a fortune, ruling them out too.

Accessibility Trees Are a Hack

Most computer use agents scrape accessibility information (like bounding boxes and text boxes) and send it to the model. But accessibility trees vary everywhere: every operating system has its own format, and different applications can structure their trees differently. Some applications always build their tree and some don't, and the latency to build one is highly variable. That makes them hard to work with, especially for an OS-agnostic tool that has to normalize whatever tree it receives. I actually built a library that unified most of these trees into a single format.

When testing Glider, one day I had the accessibility info disabled. I forgot and actually noticed no accuracy drop. But the reduced token count meant lower latency. This was discovered completely by accident. By the end of the project, I had removed the accessibility information altogether after more testing.

Relying on accessibility information is the wrong long-term path for computer use agents. These agents are supposed to be generic and work in all environments. Dependency on accessibility trees stifles this flexibility. And models are now intelligent enough to decode screenshots on their own without the extra info.

Glider Needs New Architectures

LLMs are not built for computer use agents. I learned this the hard way; you've probably heard me complain about latency enough by now. New architectures are needed to master these kinds of tasks. Fortunately, the AI community knows this too and great progress has been made so far.

One development I find especially promising is TypeSafe AI's model called Jev, the first decision model. Instead of generating text token by token, Jev returns typed decisions with calibrated probabilities, in a single, very fast call. For example, it can answer hundreds of multiple-choice questions in parallel in a tiny fraction of a second. Developers have already begun using Jev to create blazing-fast computer use agents, by asking structured decision questions: “Should I click box 1? Should I click box 2? …” The downside is that it doesn't support images yet, leaving it doomed to accessibility trees, but I expect that to change soon.

While image-based decision models would make Glider significantly more powerful, I truly think an even more advanced architecture is needed. Decision models still require text context, like needing to tell them how much progress has been made on the task so far. This keeps long-horizon tasks out of reach.

What we need is some sort of model with an internal memory like a built-in latent state, rather than having the model egest and ingest English over and over. This would also follow Rich Sutton's “The Bitter Lesson”: methods that learn their own representations tend to beat the ones we design by hand. Right now, we keep context in English, a fairly arbitrary, human-defined language. If a model instead learned a latent state that represents real-world concepts, it might be much better at remembering things. This even goes well beyond computer use agents.

Where Glider Is Now

Glider exists today, but I haven't published it. For now, I'm keeping an eye out for next-generation architectures that can carry it forward. Qwen 4 27B is already in the works, and I can't wait to try it out (especially after converting it to a decision model). But context and latency are the bottlenecks now, not intelligence. So Glider has a clear direction, and I'm not calling the project dead, but for now you'll see me working on other projects. And to those hoping for an agent that can play League of Legends for them: you'll have to wait a little longer.