diff --git a/README.md b/README.md index da1115c..fec9e84 100644 --- a/README.md +++ b/README.md @@ -1,21 +1,17 @@ -
-
-
English · 简体中文 · Español · 日本語 · 한국어 · Português (Brasil) · Français · Deutsch · Русский · العربية · Italiano
- -## Why it exists - -browser-search is a SKILL — an instruction set for AI agents like OpenCode, -Claude Code, Cursor, OpenClaw and others. It teaches your agent how to -search and browse the web using three orchestrated open source tools. - -The problem? The web is hostile to automation. Cloudflare, Akamai, DataDome -and other anti-bot systems block simple requests. Modern sites use heavy -JavaScript, lazy loading, and client-side rendering. One single solution -is not enough. - -`browser-search` orchestrates **three open source tools** into a single -search and browsing system designed for AI agents. Each tool has its role, -orchestrated by the skill with escalation logic, automatic selection, -and ready-to-use integration: - -1. **[SearXNG](https://github.com/searxng/searxng)** — metasearch engine for the search phase (multi-source, JSON) -2. **[Camofox](https://github.com/jo-inc/camofox-browser)** — browser navigable via REST API for standard sites -3. **[CloakBrowser](https://github.com/cloakhq/cloakbrowser)** — stealth browser for anti-bot protected sites - -The typical flow: the agent first searches with SearXNG, then browses the -results with Camofox (or CloakBrowser if the site is protected). - -## Benefits - -- **100% free, self-hosted, unlimited.** No API keys to buy, no - subscriptions, no rate limits. Everything runs on your machine, - Docker and npm. Unlimited usage, zero cost. - -- **Lightweight, runs anywhere.** Built and tested on a Raspberry Pi - — if it runs there, it runs everywhere. Minimal resource consumption, - no heavy infrastructure needed, runs 24/7 on low-power hardware. - -- **Search + browse in one kit.** No manual integration needed. - Searching and browsing are two distinct phases, both covered. - -- **Automatic navigation escalation.** If Camofox gets blocked by - Cloudflare/Akamai, the agent automatically switches to CloakBrowser. - -- **Smart performance.** SearXNG for the search phase (milliseconds). - Camofox and CloakBrowser are only used to browse the sites that - actually need it. - -- **Automatic agent choice.** The AI agent decides which tool to use: - SearXNG for initial search, Camofox for browsing, CloakBrowser if - the site is protected. Zero human intervention. - -- **Deep Research mode.** The skill instructs the agent to go beyond - superficial answers: explore multiple angles, cross-verify sources, - cover every aspect, and never cut corners. - -- **Fully customizable.** The SKILL.md is plain text. You can edit the - core rules, add your own, remove what you don't need. Adapt it to - your workflow, your team, your standards. - -- **Native stealth.** CloakBrowser automatically detects Cloudflare, - Akamai, DataDome, Imperva, PerimeterX, and DDoS-Guard challenges, - and waits for them to resolve before extracting content. - -- **Works with any agent.** The SKILL.md is written for OpenCode, - but the logic is identical for any AI agent. Same README, same - package.json, everything works everywhere. Just ask your agent - how to convert the skill for its environment. - -## 🏆 State of the art - -These three tools were chosen because they represent the current -state of the art available today. A skill like this is designed -to evolve: when better tools emerge, updating the SKILL.md is -all it takes to swap them in. 🔄 - -⭐ **Star the repo and follow** to stay up to date on new tools, -flow improvements, and orchestration updates over time. 🚀 - -## Architecture - -``` -┌─────────────────────────────────────────────────────────┐ -│ browser-search │ -│ │ -│ ┌──────────────┐ │ -│ │ Search │ │ -│ │ │ │ -│ │ SearXNG │ search engines → URLs │ -│ │ (Docker) │ JSON results, fast │ -│ │ :8080 │ │ -│ └──────────────┘ │ -│ │ │ -│ │ results ready → to browse │ -│ ↓ │ -│ ┌─────────────────────────────────────┐ │ -│ │ Browsing │ │ -│ │ │ │ -│ │ ┌──────────────┐ │ │ -│ │ │ Camofox │ browser + REST │ │ -│ │ │ (Docker) │ JS, click, eval │ │ -│ │ │ :9377 │ │ │ -│ │ └──────┬───────┘ │ │ -│ │ │ │ │ -│ │ │ if blocked │ │ -│ │ ↓ │ │ -│ │ ┌──────────────┐ │ │ -│ │ │ CloakBrowser │ stealth Chromium │ │ -│ │ │ (npm) │ anti-bot, proxy │ │ -│ │ └──────────────┘ │ │ -│ └─────────────────────────────────────┘ │ -└─────────────────────────────────────────────────────────┘ -``` - -## How it works - -### Phase 1 — Search with SearXNG - -Docker container on `localhost:8080`. Metasearch engine that queries -Google, Wikipedia, Bing, DuckDuckGo and many others simultaneously. -JSON output with titles, snippets, and URLs. - -**Example:** - -```bash -curl -s "http://localhost:8080/search?format=json&q=largest+llm+benchmark+2026" -``` - -The agent now has a list of URLs to visit and autonomously decides -whether to browse them with Camofox or CloakBrowser based on the site. - -### Phase 2 — Browse with Camofox - -Docker container on `localhost:9377`. Exposes a full Firefox browser -through a REST API. The agent can create tabs, navigate, click, -scroll, execute arbitrary JavaScript, and structure data. - -**Includes:** Mozilla's Readability.js for extracting clean articles, -removing nav, sidebar, and ads (~70% token savings). - -**Main commands:** - -```bash -# Create tab and navigate -curl -s -X POST "http://localhost:9377/tabs" \ - -H 'Content-Type: application/json' \ - -d '{"userId":"bot","url":"https://example.com"}' - -# Read snapshot (accessibility tree) -curl -s "http://localhost:9377/tabs/
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-