AI Integration
AI SDK
AI vision-based content matching using LLMs via Vercel AI SDK.
Overview
The @nut-tree/plugin-ai-sdk plugin enables AI vision-based content matching in nut.js using large language models via the Vercel AI SDK. Instead of pixel-based template matching, it uses multimodal LLMs to understand screen content by natural language description, both for locating UI elements and for formulating test expectations.
Content Matching
Find UI elements by describing what you see
screen.find(contentMatchingDescription("a login button"))Test Assertions
Write visual expectations in natural language
expect(screen).toShow(contentMatchingDescription("a welcome dialog"))Hosted or Self-Hosted
OpenAI, Anthropic, Ollama, and any server speaking the OpenAI Chat Completions API
useOpenAICompatibleVisionProvider()What's New in 2.0
Version 2.0 lets you run nut.js vision matching against self-hosted models, in the coordinate convention they were trained on, with confidence scores you can actually threshold. It adds three things:
- A first-class provider for OpenAI-compatible servers such as vLLM, LM Studio and llama.cpp.
- A configurable coordinate space so models report regions the way they were trained to.
- A rewritten confidence instruction that makes the plugin's threshold meaningful.
It ships together with @nut-tree/vision-evals, a benchmark harness that measures models through this exact pipeline.
Major Release
Installation
npm install @nut-tree/plugin-ai-sdkSubscription Required
Quick Reference
Provider Functions
Each helper registers a vision finder for one backend. All of them accept the shared configuration plus a model override.
useOpenAIVisionProvider
useOpenAIVisionProvider(options?)Activate OpenAI as the vision provider. Model defaults to gpt-5-mini.
useAnthropicVisionProvider
useAnthropicVisionProvider(options?)Activate Anthropic as the vision provider.
useOllamaVisionProvider
useOllamaVisionProvider(options?)Activate Ollama as the vision provider for local matching. Model defaults to llama3.2-vision:11b.
useOpenAICompatibleVisionProvider
useOpenAICompatibleVisionProvider(options)New in 2.0. Activate a self-hosted server that speaks the OpenAI Chat Completions API (vLLM, LM Studio, llama.cpp, LocalAI). Model ID and server config are required.
Query Functions
contentMatchingDescription
contentMatchingDescription(description: string)Creates a content query that describes what to look for on screen. Used with screen.find, screen.findAll, screen.waitFor, and test matchers like toShow and toMatchContentDescription.
Provider Setup
OpenAI
import { screen, contentMatchingDescription } from "@nut-tree/nut-js";
import { useOpenAIVisionProvider } from "@nut-tree/plugin-ai-sdk";
// Picks up OPENAI_API_KEY from the environment
useOpenAIVisionProvider({ model: "gpt-5-mini" });
// Find elements by description
const region = await screen.find(
contentMatchingDescription("the login button")
);API Key Required
OPENAI_API_KEY environment variable, or pass it via openai.apiKey in the configuration.Anthropic
import { screen, contentMatchingDescription } from "@nut-tree/nut-js";
import { useAnthropicVisionProvider } from "@nut-tree/plugin-ai-sdk";
// Picks up ANTHROPIC_API_KEY from the environment
useAnthropicVisionProvider({ model: "claude-sonnet-5" });
const region = await screen.find(
contentMatchingDescription("the submit button")
);API Key Required
ANTHROPIC_API_KEY environment variable, or pass it via anthropic.apiKey in the configuration.Ollama (Local)
Use Ollama for fully local AI matching without external API calls:
import { screen, contentMatchingDescription } from "@nut-tree/nut-js";
import { useOllamaVisionProvider } from "@nut-tree/plugin-ai-sdk";
useOllamaVisionProvider({
model: "llama3.2-vision:11b",
ollama: { baseURL: "http://localhost:11434/api" },
});
const region = await screen.find(
contentMatchingDescription("the search input field")
);Local Setup
ollama serve) and you have pulled a vision-capable model.OpenAI-Compatible Servers (vLLM, LM Studio, llama.cpp)
Most self-hosted inference servers speak the OpenAI Chat Completions API. The official OpenAI provider in the AI SDK, however, routes requests through OpenAI's newer Responses API, which those servers do not implement. Pointing useOpenAIVisionProvider at a local base URL therefore worked with some backends and failed with others, most visibly with vLLM.
The openai-compatible provider, new in 2.0 and built on the AI SDK's dedicated OpenAI-compatible adapter, talks Chat Completions only:
import { screen, contentMatchingDescription } from "@nut-tree/nut-js";
import { useOpenAICompatibleVisionProvider } from "@nut-tree/plugin-ai-sdk";
useOpenAICompatibleVisionProvider({
model: "Qwen/Qwen3-VL-8B-Instruct",
openaiCompatible: {
baseURL: "http://localhost:8000/v1",
},
});
const region = await screen.find(
contentMatchingDescription("the 'Sign in' button")
);Both the model ID and the server config are required, because there is no sensible default for a self-hosted setup. The server config accepts:
| Option | Meaning |
|---|---|
baseURL | Base URL of the server, including the /v1 suffix if the server expects it. Required. |
apiKey | Optional bearer token, sent as an Authorization header when set. |
name | Provider name reported to the AI SDK. Defaults to openai-compatible. |
headers | Extra HTTP headers sent with every request. |
queryParams | Extra query parameters appended to every request. |
supportsStructuredOutputs | Whether the server accepts a JSON-schema response_format. Defaults to true. |
The openai-compatible value is also accepted in the provider field of a per-request model override, so a single finder can mix hosted and local models per call.
Two details that matter for local backends
Images are sent as raw base64. The finder strips the data:image/png;base64, prefix and transmits the image as a typed file part with an explicit media type. Several compatible servers rejected the full data URL that 1.x sent.
Structured JSON output is requested explicitly. The finder always parses the model's answer against a JSON schema, so by default it asks the server for a JSON-schema response format. Backends that do not implement structured outputs can set supportsStructuredOutputs: false, in which case the plugin falls back to parsing the plain text answer.
Configuration
All provider setup functions accept the same configuration object:
import { useOpenAIVisionProvider } from "@nut-tree/plugin-ai-sdk";
useOpenAIVisionProvider({
// Model ID for this provider
model: "gpt-5-mini",
// Confidence threshold in the range 0 to 1 (default 0.8)
defaultConfidence: 0.7,
// Maximum matches to return per search
defaultMaxMatches: 5,
// Coordinate convention the model was trained on (default "pixels")
coordinateSpace: "pixels",
// Provider connection options
openai: { apiKey: process.env.OPENAI_API_KEY },
});Options Reference
model
model?: stringModel ID to send to the provider. Optional for OpenAI, Anthropic and Ollama, which have defaults. Required for OpenAI-compatible servers.
defaultConfidence
defaultConfidence?: numberDefault confidence threshold for matches in the range 0 to 1 (default 0.8). Can be overridden per search.
defaultMaxMatches
defaultMaxMatches?: numberMaximum number of matches to return per search. Can be overridden per search.
coordinateSpace
coordinateSpace?: "pixels" | "normalized"New in 2.0. Coordinate space the model reports regions in (default "pixels").
coordinateScale
coordinateScale?: numberNew in 2.0. Upper bound of the normalised coordinate range (default 1000). Only used when coordinateSpace is "normalized".
matching
matching?: VisionMatchingSettingsExplicit matcher tuning settings for confidence and max matches. When present, these values override equivalent top-level options.
openai / anthropic / ollama
{ apiKey?: string; baseURL?: string }Connection options for the hosted providers and Ollama. Ollama accepts baseURL only.
openaiCompatible
openaiCompatible?: { baseURL: string; ... }New in 2.0. Connection details for a self-hosted OpenAI-compatible server. Required when the provider is openai-compatible.
Coordinate Space
Vision models do not agree on how to express a bounding box. Some families, such as Qwen2.5-VL, are trained to answer in absolute pixels of the input image. Others, such as Gemma and Qwen3-VL, are trained on a normalised grid, typically 0 to 1000, relative to the image size. Version 1.x always asked for pixels. A model trained on a normalised grid would still comply, but its numbers were in the wrong unit, and every match landed off by a constant factor. From the outside that looks like a model that "cannot localise", when in fact it was answering a question it was never trained to answer.
Two options on the provider config select the convention:
useOpenAICompatibleVisionProvider({
model: "google/gemma-3-12b-it",
openaiCompatible: { baseURL: "http://localhost:1234/v1" },
coordinateSpace: "normalized",
coordinateScale: 1000,
});| Option | Values | Default |
|---|---|---|
coordinateSpace | "pixels" or "normalized" | "pixels" |
coordinateScale | Upper bound of the normalised range | 1000 |
In pixels mode the prompt states the exact width and height of the transmitted image. That is new in 2.0 as well: in 1.x the model had to infer the dimensions, and models that reason on a downscaled copy of the image would report coordinates for the wrong size.
In normalized mode the model is asked for values in the range 0 to coordinateScale, with left and width relative to the image width and top and height relative to the image height. The finder converts the answer back to pixels deterministically before the usual region sanitisation runs. Because the convention is relative, it is also invariant under uniform rescaling by the serving runtime.
Nothing else in the pipeline changes. Consumers still receive pixel regions, centerOf still yields a pixel click point, and confidence filtering is unaffected.
How to pick
Confidence Scores
The 1.x system prompt told the model to "be confident" and to return a confidence of 1 when it was sure. In practice models anchored on that value and reported 1.0 for nearly everything, including wrong matches. The defaultConfidence threshold, and the confidence argument on screen.find, had very little to filter on.
In 2.0 the anchor is gone. The prompt asks the model to report confidence as its honest estimate of the probability, in the range 0 to 1, that the reported content truly matches the description, and to use the full range rather than defaulting to extreme values. The plugin's existing filtering, sorting and no-match reporting are unchanged, but they now operate on a signal with spread.
Existing setups
--confidence values and reading the recall and precision columns together.Per-Request Overrides
A match request can override the model and the maximum number of matches through its providerData. The provider field accepts any of the supported backends, including openai-compatible:
const region = await screen.find(
contentMatchingDescription("the 'Export' button"), {
providerData: {
model: { provider: "openai-compatible", model: "Qwen/Qwen3-VL-8B-Instruct" },
maxMatches: 3,
},
}
);Full Option Reference
interface VercelAiSdkVisionProviderConfig {
defaultModel?: { provider: "openai" | "anthropic" | "ollama" | "openai-compatible"; model: string };
defaultConfidence?: number; // [0, 1], default 0.8
defaultMaxMatches?: number;
coordinateSpace?: "pixels" | "normalized"; // new in 2.0, default "pixels"
coordinateScale?: number; // new in 2.0, default 1000
openai?: { apiKey?: string; baseURL?: string };
anthropic?: { apiKey?: string; baseURL?: string };
ollama?: { baseURL?: string };
openaiCompatible?: { // new in 2.0
baseURL: string;
apiKey?: string;
name?: string;
headers?: Record<string, string>;
queryParams?: Record<string, string>;
supportsStructuredOutputs?: boolean;
};
}Wiring helpers: useOpenAIVisionProvider, useAnthropicVisionProvider, useOllamaVisionProvider and, new in 2.0, useOpenAICompatibleVisionProvider. All of them accept the config above plus a model override and an optional matching block for the confidence and max-matches settings.
Usage
Finding Elements
Use contentMatchingDescription with screen.find to locate UI elements by describing what they look like:
import { screen, mouse, centerOf, straightTo, Button, contentMatchingDescription } from "@nut-tree/nut-js";
import { useOpenAIVisionProvider } from "@nut-tree/plugin-ai-sdk";
useOpenAIVisionProvider({ model: "gpt-5-mini" });
// Find a specific UI element by description
const region = await screen.find(
contentMatchingDescription("a system menu called 'Navigate'"), {
confidence: 0.8,
}
);
// Interact with the found region
await mouse.move(straightTo(centerOf(region)));
await mouse.click(Button.LEFT);Waiting for Elements
Wait for content to appear on screen within a timeout:
// Wait up to 10 seconds for a dialog to appear, checking every second
const dialog = await screen.waitFor(
contentMatchingDescription("a confirmation dialog with 'Save changes?' text"),
10000,
1000
);Test Assertions
The real power of contentMatchingDescription shines in end-to-end tests, where you can formulate visual expectations in plain language:
import { screen, contentMatchingDescription } from "@nut-tree/nut-js";
import { useOpenAIVisionProvider } from "@nut-tree/plugin-ai-sdk";
useOpenAIVisionProvider({ model: "gpt-5-mini" });
// Assert that the screen shows specific content
await expect(screen).toShow(
contentMatchingDescription("a browser window with an open npmjs.org tab")
);
// Assert with a custom confidence threshold
await expect(screen).toShow(
contentMatchingDescription("a navigation bar with a 'Home' link"),
{ confidence: 0.9 }
);The toShow matcher is available through the Jest and Vitest integration matchers provided by @nut-tree/nut-js.
Full E2E Test Example
Here is a complete example combining content matching with test assertions:
import { describe, it, expect, beforeAll } from "vitest";
import { screen, mouse, keyboard, centerOf, straightTo, Button, contentMatchingDescription } from "@nut-tree/nut-js";
import { useOpenAIVisionProvider } from "@nut-tree/plugin-ai-sdk";
beforeAll(() => {
useOpenAIVisionProvider({ model: "gpt-5-mini" });
});
describe("Application navigation", () => {
it("should open the settings page", async () => {
// Find and click the settings menu
const settingsMenu = await screen.find(
contentMatchingDescription("a menu item labeled 'Settings'")
);
await mouse.move(straightTo(centerOf(settingsMenu)));
await mouse.click(Button.LEFT);
// Verify the settings page is displayed
await expect(screen).toShow(
contentMatchingDescription("a settings page with 'General' and 'Advanced' tabs")
);
});
it("should display search results", async () => {
await keyboard.type("nut.js automation");
await expect(screen).toShow(
contentMatchingDescription("a list of search results related to 'nut.js'")
);
});
});Upgrading to 2.0
- Self-hosted backend? Switch from
useOpenAIVisionProviderwith a custom base URL touseOpenAICompatibleVisionProvider. Include the/v1suffix in the base URL if your server expects it. - Backend without structured outputs? Set
supportsStructuredOutputs: falsein the server config. - Check the coordinate convention for your model family. Run the coordinate-space A/B from the eval package's example configs rather than assuming.
- Re-check your confidence threshold. Scores are no longer anchored at 1.0. Run the eval suite at a few thresholds and pick the one where recall and precision meet your needs.
- Pin the eval package to the plugin version you deploy, and re-run it after every plugin or model upgrade.
Benchmark harness
Best Practices
Writing Good Descriptions
- Be specific about what you see (e.g., "a system menu called 'Navigate'" vs "a menu")
- Include visual characteristics like color, position, or text content
- Describe the element in context (e.g., "a browser window with an open npmjs.org tab")
Performance Considerations
- AI vision matching is slower than template matching (nl-matcher) due to API latency
- Cloud providers (OpenAI, Anthropic) require internet connectivity and incur API costs
- Ollama or self-hosted OpenAI-compatible models provide local matching but require a capable GPU for good performance
- Consider using nl-matcher for speed-critical operations and AI SDK for complex visual understanding