IT Brief US - Technology news for CIOs & IT decision-makers
United States
Nvidia & Microsoft push local AI agents on Windows

Nvidia & Microsoft push local AI agents on Windows

Fri, 4th Sep 2026 (Today)
Sean Mitchell
SEAN MITCHELL Publisher

NVIDIA, Microsoft and several software partners have expanded support for running AI agents locally on NVIDIA hardware, alongside new Windows PCs in the RTX Spark line.

The announcements focus on easier installation for local agent software, faster inference in widely used open-source tools, and a new software layer called NVIDIA PAIR, which distributes inference jobs across multiple PCs on a local network.

Three agent applications are adding simplified local model setup on Windows systems with NVIDIA GPUs: Hermes Agent, OpenClaw and Perplexity Portable Computer. The changes are intended to reduce the manual configuration previously needed to choose models, match them with inference servers and tune settings for local use.

Hermes Agent, developed by Nous Research, will offer one-click local setup across RTX and DGX systems on Windows. The software is designed to detect the installed NVIDIA GPU, choose a model and configuration, and run it through an integrated version of llama.cpp with NVIDIA optimisations already applied.

Once installed, Hermes is intended to maintain context across tasks, retain information between sessions and build reusable skills over time. Linux support for the simplified setup is due later.

OpenClaw is also adding a simplified Windows setup process for local models on RTX GPUs with at least 24GB of VRAM. The project has built a large open-source following, and NVIDIA, Microsoft and OpenClaw said they worked together on the Windows application to reduce barriers for users installing local models.

Perplexity Portable Computer, which already runs locally on Linux systems including DGX Spark, is also being extended to NVIDIA RTX GPUs with at least 24GB of VRAM. Windows support is due later, bringing the same bundled approach to models, orchestration and tools to a wider PC user base.

The software lets users keep complete workflows on-device while sending selected tasks to cloud models if extra research or reasoning is needed. It asks permission before content is sent to the cloud, aiming to help users keep sensitive material on their own devices.

Inference gains

NVIDIA is also targeting performance in the open-source inference stack. New work with the llama.cpp and vLLM communities has improved throughput for local agent workloads across GeForce RTX, RTX Pro and DGX systems.

In llama.cpp, the latest updates deliver up to 1.9 times higher throughput on a GeForce RTX 5090 through kernel changes, speculative decoding improvements and faster prefill. In vLLM, throughput increases by 1.2 times on an RTX Pro 6000 Blackwell Workstation Edition and by as much as 1.4 times on two DGX Spark clusters.

Those gains are available directly in the inference backends and through applications including LM Studio and Ollama. The focus on llama.cpp and vLLM reflects their growing use among developers and enthusiasts running language models locally rather than in hosted environments.

Using spare PCs

The new NVIDIA PAIR software is aimed at users with more than one machine on the same network. PAIR, short for Personal AI Router, automatically identifies compatible PCs and routes separate inference requests to whichever system has available capacity.

The software works with Ollama and LM Studio and is available in beta for Windows, macOS and Linux. Supported hardware includes GeForce RTX 20 Series GPUs and newer, RTX Pro workstation GPUs based on Turing architecture and later, DGX Spark systems, and Apple M4 or newer silicon.

NVIDIA has positioned the tool as a way to use idle computing resources in homes and small offices where multiple PCs sit underused for parts of the day. In practice, an agent breaking a task into smaller jobs could spread them across several machines instead of waiting for one GPU to process each request in sequence.

RTX Spark rollout

The local AI push is tied to NVIDIA's forthcoming RTX Spark Windows PCs, due to arrive in October through hardware partners including Lenovo and Acer. Compact desktop and laptop designs have already been added to the line-up scheduled for release.

Lenovo has announced the Yoga Pro 9n and Yoga 9n 2-in-1, while Acer has shown a compact desktop concept. NVIDIA said the systems are built around a Blackwell GPU and Grace CPU design intended to support creators, gamers and local AI agents on the same machine.

NVIDIA also said game publishers including Electronic Arts, Embark and Ubisoft are among those bringing titles to RTX Spark systems, adding to a list disclosed earlier.

Creative software

CyberLink is among the software companies lining up support for RTX Spark. Its PhotoDirector AI PC Mode will integrate image diffusion models into PhotoDirector 365 for tasks including image editing, enhancement, object removal and background replacement, with users able to choose between local and cloud processing.

On NVIDIA GPUs, the software uses TensorRT-RTX and FP8 for local AI processing. The addition underscores NVIDIA's broader effort to link local AI use not just to coding and agent software, but also to consumer creative applications.

The broader backdrop is a growing industry push to move more generative AI workloads from cloud services onto local devices, particularly for users concerned about privacy, latency and ongoing usage costs. NVIDIA's latest moves show it is trying to strengthen the software layer around its GPUs as competition grows over where AI workloads run.

More than half of US households have two or more PCs, according to NVIDIA, and much of that computing power sits idle throughout the day.