Veda: a from-scratch Rust operating system with a voice AI agent built in

Veda: a from-scratch Rust operating system with a voice AI agent built in

Veda is an operating system with an AI agent built into it, according to its GitHub README. The project's framing is that the agent is not an assistant bolted onto a desktop but a part of the system that every application is made to work with. The user talks to it the way they would talk to someone in the room, for example "Hey Veda, switch my wallpaper to the dunes", and it does the work through the same operations as the keyboard and mouse, in any application, asking before anything that is hard to undo.

The README says everything is written from scratch in Rust and lives in the one repository: the UEFI bootloader, a capability-based microkernel, the drivers, the system services, the window system, the toolkit and the applications. A few mature crates supply the TCP/IP engine, TLS and cryptography (smoltcp, rustls, RustCrypto). Veda runs on x86-64 PCs and in QEMU.

The argument behind the design is that operating systems were built for a person at a keyboard. An assistant added later usually has to read the screen, imitate clicks or make do with the few applications that offer an API, and the system has no say in what it does. Veda turns this around: the agent is a first-class user of the system, and the system is built so that its work is reliable, visible and safe.

In practice that means several things. The agent is a system service, started at boot and restarted if it fails, shown as a small ring on the taskbar and asleep until called by name, by a click on the ring or with Super+Space. Applications describe their actions with typed parameters and a risk level, report what they show as structured state, and carry out actions through the same code as the keyboard and mouse, so the window shows the result as if the user had done it. There is no screen scraping and no simulated clicks. The built-in applications offer more than 70 actions between them, from writing in a document to starting a race.

Consent is the system's job. Every action is routine, sensitive or destructive. Routine actions just happen; the others wait for the user's OK. The agent service enforces this, not the language model: the call is held until the user answers, only the desktop shell can answer, and nothing the agent reads in a document or an application can approve anything. A held call is shown in the agent's window, or as a notification while the window is closed, saying exactly what will happen and to what. If the user declines or lets it expire, it does not run. Sensitive actions can be allowed for good; destructive ones are asked about every time.

Every conversation starts with context: the time, the windows on the screen, every application's actions, what the user has said about themselves and how the last conversation ended. While it works the agent reads the applications' state. Its memory holds what the user tells it plainly, such as a name, the people in their life and how they like things done, never its own guesses. That memory stays on the user's computer, and Settings shows all of it and can forget any of it. The agent's memory and key live in a private directory that only the agent can open.

Voice comes first. Conversations are real time: interrupt the agent and it stops and listens. The audio system has echo cancellation and an echo gate so the agent never hears itself, and other sound is turned down while it talks. While the agent sleeps, a voice detector on the computer listens for speech, so that nothing leaves the computer while nobody speaks. A sentence containing the agent's name wakes it, and waking opens a Deepgram Voice Agent session with the agent's instructions, the context and its functions. Speech recognition, the language model and the voice come from Deepgram's Voice Agent platform, with the audio opted out of Deepgram's model improvement. The model, the voice and even the agent's name are chosen in Settings. The agent's own functions cover windows, files, volume, Wi-Fi, wallpaper, notifications, timers and reminders, memory, the system and its running programs, and through use_app and read_app it reaches every application's actions and state.

Underneath is a microkernel, vkernel, that does only isolation, scheduling, memory and IPC: capability handles with rights, channels that carry handles, VMOs, events, futexes and interrupt objects, with SMP, x2APIC, tickless timers and XSAVE. Drivers, the file system, the window system and the agent are user-space processes supervised by init; if the window system crashes it is restarted and the desktop comes back on its own.

The desktop has a compositing window manager with decorations, shadows, animations, snapping and an Alt+Tab switcher with live thumbnails. It composes on the GPU with OpenGL ES on PCs with Intel graphics and on the processor elsewhere. The shell has wallpaper, desktop icons, a taskbar, a searchable start menu, a calendar and notifications. The vuitoolkit provides widgets, menus, dialogs, vector icons and text in the Inter and JetBrains Mono fonts. Built-in applications are Text Editor, Photos, Music, Files, Terminal, Task Manager, Settings and About, plus two 3D games, Velocity (racing) and Starfall (a space shooter).

Networking is a user-space service with IPv4 and IPv6, DHCP, DNS, routing, TCP, UDP and ICMP, plus a Wi-Fi service that joins WPA2, WPA3 and open networks. Applications, the agent among them, get TLS 1.3 and 1.2 from vtls, with certificates checked against the Mozilla roots. Storage covers virtio-blk, AHCI (SATA) and NVMe drivers and a file system service with crash-safe snapshots.

Veda does not write every driver itself. It drives its disks itself (and sound, for now) and takes the rest from Linux, which it runs in a virtual machine, the driver VM, on a hypervisor of its own (VMX with EPT). Devices go to that VM whole, with DMA confined to its memory and interrupts remapped by Veda's IOMMU driver. Linux's drivers there serve keyboards, mice and tablets, USB, wired and Wi-Fi networks, displays and GPUs, and when the VM fails Veda resets its devices and starts it again.

For graphics, Veda offers OpenGL ES 3.0 with GLSL ES 1.00 and 3.00, rendered on the GPU through the driver VM or with a multi-threaded software renderer where there is no GPU. A demo called Prism, with a reflective knot, shadow-mapped crystals and a spark fountain simulated with transform feedback, runs at 75 frames a second under QEMU. Users can also write C programs in the Terminal or Text Editor and compile them with GCC 16 inside Veda; they are static ELF executables on the musl C library, whose system calls go to vposix, a POSIX layer written in Rust.

Building it takes an x64 PC with Linux, stable Rust, QEMU with the OVMF UEFI firmware and access to /dev/kvm. Using the agent needs a Deepgram API key entered in Veda's Settings. One command, cargo xtask run, builds every component, writes a disk image at target/veda/veda.img and boots it in a QEMU window; the first build also builds the driver VM's Linux and takes a while. The README's screenshots were taken in QEMU with the agent talking to a stand-in for Deepgram.

Key facts

  • Veda is an operating system written from scratch in Rust (UEFI bootloader, capability-based microkernel, drivers, services, window system, toolkit, applications), with a few crates for TCP/IP, TLS and cryptography; it runs on x86-64 PCs and in QEMU.
  • A voice AI agent runs as a system service and drives applications through typed actions with risk levels (routine, sensitive, destructive), not screen scraping or simulated clicks; the built-in apps expose more than 70 actions.
  • Consent is enforced by the agent service rather than the language model: held calls wait until the user answers, and only the desktop shell can answer.
  • Speech recognition, the language model and the voice come from Deepgram's Voice Agent platform, so the agent needs a Deepgram API key set in Settings.
  • Most hardware drivers come from Linux running in a virtual machine on Veda's own hypervisor; Veda drives disks itself and sound for now.

Why it matters

Veda is a concrete take on a question many agent products dodge: what if the operating system itself were designed for an AI agent as a user? The README argues that assistants added to a desktop have to read the screen or imitate clicks, and that the system has no say in what they do. Veda instead makes applications declare their actions with typed parameters and risk levels, and makes the system, not the model, hold risky calls until the user answers. Whether that holds up in practice is not shown in the README, but the design is specific and inspectable.

Who it affects

Mainly people who build or study operating systems, and developers thinking about how agents should be given control of applications and how consent should be enforced. Trying it means having an x86-64 PC or QEMU, and a Deepgram account for the agent. The README names no end-user audience beyond that.

How to use it

The README says to build on an x64 PC with Linux, with stable Rust, QEMU and the OVMF firmware (on Debian and Ubuntu: qemu-system-x86, qemu-system-gui, qemu-system-modules-opengl and ovmf), and access to /dev/kvm. Check the environment with cargo xtask doctor, then run cargo xtask run, which builds every component, writes target/veda/veda.img and boots it in QEMU. Options include --resolution 1920x1080, --smp 4 and --memory 2048. For the agent, enter a Deepgram API key in Settings, then wake it by name, with a click on the ring or with Super+Space. C and GCC in the image are optional and need the distribution's build tools; the driver VM needs the Linux kernel's build tools, a cross compiler and Mesa, plus KVM nested virtualization (kvm_intel nested=1) under QEMU.

How solid is it

Everything here is the project's own description in its README; there are no independent tests or reviews. The visible text names no author, organisation or team behind the project. No release date, version number or licence is given in the visible text, and it offers no benchmarks, reliability figures, security audit or user-study results for the agent. The one performance figure is the Prism 3D demo at 75 frames a second under QEMU. The README's screenshots were taken with a stand-in for Deepgram, not the live service.

Risks and caveats

The agent depends on Deepgram's cloud platform for speech recognition, the language model and the voice, and the text does not say whether the agent can work offline. The README says nothing leaves the computer while nobody speaks, because a local detector listens while the agent sleeps. The claim that the system is written from scratch has a limit: a few third-party crates supply TCP/IP, TLS and cryptography, and Linux drivers run in a virtual machine for most devices. Hardware is limited to x86-64 PCs and QEMU, GPU composition is described for Intel graphics, and the driver VM needs nested virtualization under QEMU. The first build also takes a while. Sensitive actions can be allowed permanently, so the consent model is only as strict as the user's choices.

“Hey Veda, switch my wallpaper to the dunes”

— Example voice command from the Veda README