8 Best Tarsier Alternatives in 2026 (Open Source)

Tarsier — Vision utilities for web interaction agents 👀. vs vision-language models for web tasks: OCR-to-text conversion enables text-only LLMs to outperform multimodal models by 10-20% on web interaction benchmarks — more accurate and cheaper than GPT-4V

These 8 open-source tools do the same job. They are ordered by how closely they match Tarsier, with live GitHub data so you can see which projects are actively maintained.

ToolGitHub starsStars / 30dLast commit
Tarsier(original)1.8k+12024-10-01
Browser-Use116.8k+5,1502026-09-26
Skyvern23.1k+3422026-09-30
LaVague6.4k+122025-01-21
Notte2.0k+132026-09-30
Taxy AI1.3k+12025-01-15
Self-Operating Computer10.3k+132025-09-19
VisionAgent5.3k+52025-08-19
vimGPT2.6k+-32024-09-25
  1. 1. Browser-Use

    🌐 Make websites accessible for AI agents. Automate tasks online with ease.

    What sets it apart: Unlike Selenium/Playwright (code-based browser automation), Browser Use enables natural language browser control with vision AI — the first open-source library purpose-built for LLM-driven web agent tasks

    Best for: Automating repetitive web tasks like form filling, data extraction, and e-commerce workflows; QA teams needing AI-driven browser testing without writing Selenium/Playwright scripts

  2. 2. Skyvern

    Automate browser based workflows with AI

    What sets it apart: vs Browser Use: Playwright-native SDK with AI fallback mode (selector-first, AI-second) rather than pure vision approach; vs Puppeteer/Playwright: adds AI understanding so automations survive layout changes

    Best for: Automating workflows on websites that change frequently; Non-technical users building browser automations; Cross-site data extraction without custom scrapers

  3. 3. LaVague

    Large Action Model framework to develop AI Web Agents

    What sets it apart: vs Playwright/Selenium scripts: natural language objective → autonomous browser action via World Model + Action Engine architecture, no manual selector writing needed

    Best for: Automating complex web workflows via natural language; QA teams building browser-based test automation with AI

  4. 4. Notte

    🌸 Best framework to build web agents, and deploy serverless web automation functions on reliable browser infra.

    What sets it apart: vs Browser-Use/Convergence: 2x faster task completion (47s vs 113s), 96.6% reliability, and hybrid scripting+AI approach that cuts costs 50%+ while maintaining accuracy

    Best for: Building reliable web automation agents at scale; Scraping and structured data extraction from websites; Enterprise web workflows with credential management

  5. 5. Taxy AI

    Automate your browser with GPT-4

    What sets it apart: Chrome extension using GPT-4 to parse simplified DOM and execute click/setValue actions for natural language browser automation

    Best for: ai-browser-automation-research; repetitive-browser-task-automation; web-interaction-prototyping

  6. 6. Self-Operating Computer

    A framework to enable multimodal models to operate a computer.

    What sets it apart: vs Anthropic Computer Use / Browser Use: one of the first open-source frameworks for full computer-use — multimodal models see the screen and execute mouse/keyboard actions across any application, not just browsers

    Best for: Automating computer tasks requiring visual understanding; Researching multimodal agent computer interaction; Cross-application workflow automation via screen recognition

  7. 7. VisionAgent

    This tool has been deprecated. Use Agentic Document Extraction instead.

    What sets it apart: Agentic visual AI that takes image/video prompts and automatically selects vision models to output runnable code for visual AI apps in minutes

    Best for: rapid-visual-ai-prototyping; automated-vision-code-generation; image-and-video-analysis

  8. 8. vimGPT

    Browse the web with GPT-4V and Vimium

    What sets it apart: vs DOM-based web agents: uses Vimium keyboard commands and pure vision (GPT-4V screenshots) for web interaction — no DOM parsing required, enabling navigation of any visual web content

    Best for: Research into vision-based web browsing agents; Web research automation using visual understanding; Exploring multimodal AI interaction patterns