// agent

Semantic Search

by ozzo · Jul 10, 2026 Public

Choose how to run this agent

⚡ Local runs on your GPU. For usable speed it needs a WebGPU-capable browser — Chrome or Edge on a machine with a graphics card, or an Apple Silicon Mac. Without a supported GPU, pick OpenAI or Anthropic above instead. Check your machine
Try it now

Requires an API key and an AgentOp account.

84 downloads
0 forks
0.0 rating

Description

Search your own notes or documents by meaning, not keywords — instant results with no LLM download.

What this agent can do

Semantic Search is built from the Semantic Search template. Runs fully on your own device: llama.cpp compiled to WebAssembly, GPU-accelerated through WebGPU, with no API key and no server. After the one-time model download it works offline. Can run on OpenAI models with your own API key, encrypted in your browser. Can run on Anthropic Claude models with your own API key, encrypted in your browser.

Runs in the browser

  • Document Q&A (local RAG). Answers from documents you drop in: the text is chunked and indexed locally in the browser (IndexedDB), the relevant passages are retrieved for each question, and nothing is uploaded.

Source Code

agent.py
# How many passages to return per search.
SEARCH_TOP_K = [[[SEARCH_TOP_K|5]]]


async def _ingest(source, text):
    """(JS-callable) Index text as individual passages, one per non-empty line.

    Per-line indexing keeps each note / FAQ entry / list item its own searchable
    unit (rather than merging a short blob into one chunk), so ranking is
    meaningful. Underscore-prefixed so it is never exposed to the LLM as a tool.
    Returns the number of passages stored, as a string (for the UI).
    """
    passages = [line.strip() for line in text.splitlines() if line.strip()]
    total = 0
    for passage in passages:
        total += await agentop_rag.add_document(source, passage)
    return str(total)


async def process_user_query(query):
    """Return the most semantically similar passages — no LLM generation.

    Retrieval is deterministic and runs entirely on the embedding model, so this
    template never loads a chat model.
    """
    hits = await agentop_rag.search(query, SEARCH_TOP_K)
    if not hits:
        return "No matches yet — index some text first, then search."
    lines = []
    for i, h in enumerate(hits):
        score = round(float(h["score"]) * 100)
        lines.append(f"[{i + 1}] {score}% match · {h['source']}\n{h['text']}")
    return "\n\n".join(lines)

More by ozzo

Receipt & Invoice Extractor

Based on the Receipt & Invoice Extractor template.

New Hire Handbook Q&A

Based on the New Hire Handbook Q&A template.

Contract Plain-Language Explainer

Upload a contract and get it explained in plain language — obligations, fees, deadlines, exit claus…

WhatsApp Sales Copilot

Turn a raw WhatsApp chat export into a mini CRM — typed quotes, bookings, payments and boarding pas…

Private Quote & Material Estimator

Drop competing contractor quotes and compare them side by side — totals, inclusions, exclusions — t…

Messy Itinerary Travel Planner

Drop your messy pile of booking PDFs and trip notes and get a clean day-by-day itinerary — plus war…