AI Integration

AI SDK

AI vision-based content matching using LLMs via Vercel AI SDK.

Overview

The @nut-tree/plugin-ai-sdk plugin enables AI vision-based content matching in nut.js using large language models via the Vercel AI SDK. Instead of pixel-based template matching, it uses multimodal LLMs to understand screen content by natural language description, both for locating UI elements and for formulating test expectations.

Content Matching

Find UI elements by describing what you see

screen.find(contentMatchingDescription("a login button"))

Test Assertions

Write visual expectations in natural language

expect(screen).toShow(contentMatchingDescription("a welcome dialog"))

Hosted or Self-Hosted

OpenAI, Anthropic, Ollama, and any server speaking the OpenAI Chat Completions API

useOpenAICompatibleVisionProvider()

What's New in 2.0

Version 2.0 lets you run nut.js vision matching against self-hosted models, in the coordinate convention they were trained on, with confidence scores you can actually threshold. It adds three things:

It ships together with @nut-tree/vision-evals, a benchmark harness that measures models through this exact pipeline.

Major Release

Two of these changes alter what the model is asked and how its answer is interpreted. The same model can behave differently on 2.0 than on 1.x, and existing setups should re-check their confidence thresholds. See the upgrade checklist.

Installation

bash
npm install @nut-tree/plugin-ai-sdk

Subscription Required

This package is included in Solo and Team subscription plans.

Quick Reference

Provider Functions

Each helper registers a vision finder for one backend. All of them accept the shared configuration plus a model override.

useOpenAIVisionProvider

useOpenAIVisionProvider(options?)
VercelAiSdkVisionFinder

Activate OpenAI as the vision provider. Model defaults to gpt-5-mini.

useAnthropicVisionProvider

useAnthropicVisionProvider(options?)
VercelAiSdkVisionFinder

Activate Anthropic as the vision provider.

useOllamaVisionProvider

useOllamaVisionProvider(options?)
VercelAiSdkVisionFinder

Activate Ollama as the vision provider for local matching. Model defaults to llama3.2-vision:11b.

useOpenAICompatibleVisionProvider

useOpenAICompatibleVisionProvider(options)
VercelAiSdkVisionFinder

New in 2.0. Activate a self-hosted server that speaks the OpenAI Chat Completions API (vLLM, LM Studio, llama.cpp, LocalAI). Model ID and server config are required.

Query Functions

contentMatchingDescription

contentMatchingDescription(description: string)
ContentQuery

Creates a content query that describes what to look for on screen. Used with screen.find, screen.findAll, screen.waitFor, and test matchers like toShow and toMatchContentDescription.


Provider Setup

OpenAI

typescript
import { screen, contentMatchingDescription } from "@nut-tree/nut-js";
import { useOpenAIVisionProvider } from "@nut-tree/plugin-ai-sdk";

// Picks up OPENAI_API_KEY from the environment
useOpenAIVisionProvider({ model: "gpt-5-mini" });

// Find elements by description
const region = await screen.find(
    contentMatchingDescription("the login button")
);

API Key Required

Set the OPENAI_API_KEY environment variable, or pass it via openai.apiKey in the configuration.

Anthropic

typescript
import { screen, contentMatchingDescription } from "@nut-tree/nut-js";
import { useAnthropicVisionProvider } from "@nut-tree/plugin-ai-sdk";

// Picks up ANTHROPIC_API_KEY from the environment
useAnthropicVisionProvider({ model: "claude-sonnet-5" });

const region = await screen.find(
    contentMatchingDescription("the submit button")
);

API Key Required

Set the ANTHROPIC_API_KEY environment variable, or pass it via anthropic.apiKey in the configuration.

Ollama (Local)

Use Ollama for fully local AI matching without external API calls:

typescript
import { screen, contentMatchingDescription } from "@nut-tree/nut-js";
import { useOllamaVisionProvider } from "@nut-tree/plugin-ai-sdk";

useOllamaVisionProvider({
    model: "llama3.2-vision:11b",
    ollama: { baseURL: "http://localhost:11434/api" },
});

const region = await screen.find(
    contentMatchingDescription("the search input field")
);

Local Setup

Make sure Ollama is running locally (ollama serve) and you have pulled a vision-capable model.

OpenAI-Compatible Servers (vLLM, LM Studio, llama.cpp)

Most self-hosted inference servers speak the OpenAI Chat Completions API. The official OpenAI provider in the AI SDK, however, routes requests through OpenAI's newer Responses API, which those servers do not implement. Pointing useOpenAIVisionProvider at a local base URL therefore worked with some backends and failed with others, most visibly with vLLM.

The openai-compatible provider, new in 2.0 and built on the AI SDK's dedicated OpenAI-compatible adapter, talks Chat Completions only:

typescript
import { screen, contentMatchingDescription } from "@nut-tree/nut-js";
import { useOpenAICompatibleVisionProvider } from "@nut-tree/plugin-ai-sdk";

useOpenAICompatibleVisionProvider({
    model: "Qwen/Qwen3-VL-8B-Instruct",
    openaiCompatible: {
        baseURL: "http://localhost:8000/v1",
    },
});

const region = await screen.find(
    contentMatchingDescription("the 'Sign in' button")
);

Both the model ID and the server config are required, because there is no sensible default for a self-hosted setup. The server config accepts:

OptionMeaning
baseURLBase URL of the server, including the /v1 suffix if the server expects it. Required.
apiKeyOptional bearer token, sent as an Authorization header when set.
nameProvider name reported to the AI SDK. Defaults to openai-compatible.
headersExtra HTTP headers sent with every request.
queryParamsExtra query parameters appended to every request.
supportsStructuredOutputsWhether the server accepts a JSON-schema response_format. Defaults to true.

The openai-compatible value is also accepted in the provider field of a per-request model override, so a single finder can mix hosted and local models per call.

Two details that matter for local backends

Images are sent as raw base64. The finder strips the data:image/png;base64, prefix and transmits the image as a typed file part with an explicit media type. Several compatible servers rejected the full data URL that 1.x sent.

Structured JSON output is requested explicitly. The finder always parses the model's answer against a JSON schema, so by default it asks the server for a JSON-schema response format. Backends that do not implement structured outputs can set supportsStructuredOutputs: false, in which case the plugin falls back to parsing the plain text answer.


Configuration

All provider setup functions accept the same configuration object:

typescript
import { useOpenAIVisionProvider } from "@nut-tree/plugin-ai-sdk";

useOpenAIVisionProvider({
    // Model ID for this provider
    model: "gpt-5-mini",

    // Confidence threshold in the range 0 to 1 (default 0.8)
    defaultConfidence: 0.7,

    // Maximum matches to return per search
    defaultMaxMatches: 5,

    // Coordinate convention the model was trained on (default "pixels")
    coordinateSpace: "pixels",

    // Provider connection options
    openai: { apiKey: process.env.OPENAI_API_KEY },
});

Options Reference

model

model?: string
optional

Model ID to send to the provider. Optional for OpenAI, Anthropic and Ollama, which have defaults. Required for OpenAI-compatible servers.

defaultConfidence

defaultConfidence?: number
optional

Default confidence threshold for matches in the range 0 to 1 (default 0.8). Can be overridden per search.

defaultMaxMatches

defaultMaxMatches?: number
optional

Maximum number of matches to return per search. Can be overridden per search.

coordinateSpace

coordinateSpace?: "pixels" | "normalized"
optional

New in 2.0. Coordinate space the model reports regions in (default "pixels").

coordinateScale

coordinateScale?: number
optional

New in 2.0. Upper bound of the normalised coordinate range (default 1000). Only used when coordinateSpace is "normalized".

matching

matching?: VisionMatchingSettings
optional

Explicit matcher tuning settings for confidence and max matches. When present, these values override equivalent top-level options.

openai / anthropic / ollama

{ apiKey?: string; baseURL?: string }
optional

Connection options for the hosted providers and Ollama. Ollama accepts baseURL only.

openaiCompatible

openaiCompatible?: { baseURL: string; ... }
optional

New in 2.0. Connection details for a self-hosted OpenAI-compatible server. Required when the provider is openai-compatible.

Coordinate Space

Vision models do not agree on how to express a bounding box. Some families, such as Qwen2.5-VL, are trained to answer in absolute pixels of the input image. Others, such as Gemma and Qwen3-VL, are trained on a normalised grid, typically 0 to 1000, relative to the image size. Version 1.x always asked for pixels. A model trained on a normalised grid would still comply, but its numbers were in the wrong unit, and every match landed off by a constant factor. From the outside that looks like a model that "cannot localise", when in fact it was answering a question it was never trained to answer.

Two options on the provider config select the convention:

typescript
useOpenAICompatibleVisionProvider({
    model: "google/gemma-3-12b-it",
    openaiCompatible: { baseURL: "http://localhost:1234/v1" },
    coordinateSpace: "normalized",
    coordinateScale: 1000,
});
OptionValuesDefault
coordinateSpace"pixels" or "normalized""pixels"
coordinateScaleUpper bound of the normalised range1000

In pixels mode the prompt states the exact width and height of the transmitted image. That is new in 2.0 as well: in 1.x the model had to infer the dimensions, and models that reason on a downscaled copy of the image would report coordinates for the wrong size.

In normalized mode the model is asked for values in the range 0 to coordinateScale, with left and width relative to the image width and top and height relative to the image height. The finder converts the answer back to pixels deterministically before the usual region sanitisation runs. Because the convention is relative, it is also invariant under uniform rescaling by the serving runtime.

Nothing else in the pipeline changes. Consumers still receive pixel regions, centerOf still yields a pixel click point, and confidence filtering is unaffected.

How to pick

Do not guess. The difference between the right and the wrong convention for a given model is dramatic, and it is precisely the kind of thing vision-evals was built to measure: a compare config can list the same model twice with both settings and rank them side by side.

Confidence Scores

The 1.x system prompt told the model to "be confident" and to return a confidence of 1 when it was sure. In practice models anchored on that value and reported 1.0 for nearly everything, including wrong matches. The defaultConfidence threshold, and the confidence argument on screen.find, had very little to filter on.

In 2.0 the anchor is gone. The prompt asks the model to report confidence as its honest estimate of the probability, in the range 0 to 1, that the reported content truly matches the description, and to use the full range rather than defaulting to extreme values. The plugin's existing filtering, sorting and no-match reporting are unchanged, but they now operate on a signal with spread.

Existing setups

A model that reported 1.0 on 1.x will report lower numbers on 2.0 for the same, correct match. Setups that relied on a high threshold, or on the default of 0.8 together with a model that happens to be conservative, may see matches that used to pass now fall below the bar. That is the threshold finally doing its job, but thresholds should be re-checked after upgrading, ideally by running the eval suite with different --confidence values and reading the recall and precision columns together.

Per-Request Overrides

A match request can override the model and the maximum number of matches through its providerData. The provider field accepts any of the supported backends, including openai-compatible:

typescript
const region = await screen.find(
    contentMatchingDescription("the 'Export' button"), {
        providerData: {
            model: { provider: "openai-compatible", model: "Qwen/Qwen3-VL-8B-Instruct" },
            maxMatches: 3,
        },
    }
);

Full Option Reference

typescript
interface VercelAiSdkVisionProviderConfig {
  defaultModel?: { provider: "openai" | "anthropic" | "ollama" | "openai-compatible"; model: string };
  defaultConfidence?: number;        // [0, 1], default 0.8
  defaultMaxMatches?: number;
  coordinateSpace?: "pixels" | "normalized";   // new in 2.0, default "pixels"
  coordinateScale?: number;                    // new in 2.0, default 1000
  openai?: { apiKey?: string; baseURL?: string };
  anthropic?: { apiKey?: string; baseURL?: string };
  ollama?: { baseURL?: string };
  openaiCompatible?: {                         // new in 2.0
    baseURL: string;
    apiKey?: string;
    name?: string;
    headers?: Record<string, string>;
    queryParams?: Record<string, string>;
    supportsStructuredOutputs?: boolean;
  };
}

Wiring helpers: useOpenAIVisionProvider, useAnthropicVisionProvider, useOllamaVisionProvider and, new in 2.0, useOpenAICompatibleVisionProvider. All of them accept the config above plus a model override and an optional matching block for the confidence and max-matches settings.


Usage

Finding Elements

Use contentMatchingDescription with screen.find to locate UI elements by describing what they look like:

typescript
import { screen, mouse, centerOf, straightTo, Button, contentMatchingDescription } from "@nut-tree/nut-js";
import { useOpenAIVisionProvider } from "@nut-tree/plugin-ai-sdk";

useOpenAIVisionProvider({ model: "gpt-5-mini" });

// Find a specific UI element by description
const region = await screen.find(
    contentMatchingDescription("a system menu called 'Navigate'"), {
        confidence: 0.8,
    }
);

// Interact with the found region
await mouse.move(straightTo(centerOf(region)));
await mouse.click(Button.LEFT);

Waiting for Elements

Wait for content to appear on screen within a timeout:

typescript
// Wait up to 10 seconds for a dialog to appear, checking every second
const dialog = await screen.waitFor(
    contentMatchingDescription("a confirmation dialog with 'Save changes?' text"),
    10000,
    1000
);

Test Assertions

The real power of contentMatchingDescription shines in end-to-end tests, where you can formulate visual expectations in plain language:

typescript
import { screen, contentMatchingDescription } from "@nut-tree/nut-js";
import { useOpenAIVisionProvider } from "@nut-tree/plugin-ai-sdk";

useOpenAIVisionProvider({ model: "gpt-5-mini" });

// Assert that the screen shows specific content
await expect(screen).toShow(
    contentMatchingDescription("a browser window with an open npmjs.org tab")
);

// Assert with a custom confidence threshold
await expect(screen).toShow(
    contentMatchingDescription("a navigation bar with a 'Home' link"),
    { confidence: 0.9 }
);

The toShow matcher is available through the Jest and Vitest integration matchers provided by @nut-tree/nut-js.

Full E2E Test Example

Here is a complete example combining content matching with test assertions:

typescript
import { describe, it, expect, beforeAll } from "vitest";
import { screen, mouse, keyboard, centerOf, straightTo, Button, contentMatchingDescription } from "@nut-tree/nut-js";
import { useOpenAIVisionProvider } from "@nut-tree/plugin-ai-sdk";

beforeAll(() => {
    useOpenAIVisionProvider({ model: "gpt-5-mini" });
});

describe("Application navigation", () => {
    it("should open the settings page", async () => {
        // Find and click the settings menu
        const settingsMenu = await screen.find(
            contentMatchingDescription("a menu item labeled 'Settings'")
        );
        await mouse.move(straightTo(centerOf(settingsMenu)));
        await mouse.click(Button.LEFT);

        // Verify the settings page is displayed
        await expect(screen).toShow(
            contentMatchingDescription("a settings page with 'General' and 'Advanced' tabs")
        );
    });

    it("should display search results", async () => {
        await keyboard.type("nut.js automation");

        await expect(screen).toShow(
            contentMatchingDescription("a list of search results related to 'nut.js'")
        );
    });
});

Upgrading to 2.0

  1. Self-hosted backend? Switch from useOpenAIVisionProvider with a custom base URL to useOpenAICompatibleVisionProvider. Include the /v1 suffix in the base URL if your server expects it.
  2. Backend without structured outputs? Set supportsStructuredOutputs: false in the server config.
  3. Check the coordinate convention for your model family. Run the coordinate-space A/B from the eval package's example configs rather than assuming.
  4. Re-check your confidence threshold. Scores are no longer anchored at 1.0. Run the eval suite at a few thresholds and pick the one where recall and precision meet your needs.
  5. Pin the eval package to the plugin version you deploy, and re-run it after every plugin or model upgrade.

Benchmark harness

@nut-tree/vision-evals is released in lockstep with the plugin. Equal version numbers belong together, and installing the eval package pulls in the matching plugin build. It runs a suite of UI screenshots through the plugin's real system prompt, output schema, confidence filter and coordinate handling, and reports whether each match is safe to click and whether the model stays silent on content that is not there.

Best Practices

Writing Good Descriptions

  • Be specific about what you see (e.g., "a system menu called 'Navigate'" vs "a menu")
  • Include visual characteristics like color, position, or text content
  • Describe the element in context (e.g., "a browser window with an open npmjs.org tab")

Performance Considerations

  • AI vision matching is slower than template matching (nl-matcher) due to API latency
  • Cloud providers (OpenAI, Anthropic) require internet connectivity and incur API costs
  • Ollama or self-hosted OpenAI-compatible models provide local matching but require a capable GPU for good performance
  • Consider using nl-matcher for speed-critical operations and AI SDK for complex visual understanding

Was this page helpful?