Answers grounded in sources you can inspect.

Turn your web estate, SharePoint, cloud drives and documents into knowledge agents can use — kept current on a schedule, searched with hybrid retrieval and cited with relevance scores on every answer.

5synced source types
6upload file types
Hybridvector + full-text
Per pageusage tracking
Website crawling

Your web estate, turned into agent knowledge.

Point Gecko at a site and it crawls, extracts and embeds the content — including modern JavaScript-rendered pages and the PDFs they link to — then keeps it current on a schedule.

  • Two crawler engines — fast HTML parsing, or full JavaScript rendering for modern sites
  • Precision controls — sitemap discovery, include patterns, exclude globs, CSS-selector exclusions, specific URL lists and page limits
  • Scheduled re-crawls — plus on-demand, cancellable and forced re-scrapes, with near-duplicate URLs removed
  • Jurisdiction-aware — choose a UK egress proxy per site, and an identifiable GeckoBot user agent
Knowledge sources

The knowledge that never made it to the website.

Policies in SharePoint, handbooks in Google Drive, procedures in OneDrive — sync the folders that matter and keep them current automatically.

SharePoint

Crawl sites, lists and pages through Microsoft Graph — extracting Word, Excel and PDF content — or sync document folders.

Google Drive

Sync folders including Google Docs, Sheets and Slides alongside PDFs, Word and text files.

OneDrive, Dropbox & S3

Bring in document folders from the storage your teams already use — including S3-compatible storage.

Set it, schedule it, see its status.

Choose a folder, include sub-folders if you want them, and set a sync schedule. Each source shows whether it is idle, pending, syncing, succeeded or failed.

  • Recursive folder sync on a cron schedule
  • PDF, Word, CSV, JSON, Markdown and text, plus Google Docs, Sheets and Slides
  • Lands in a knowledge folder you control, so access stays scoped
Knowledge base

Curate the answers you want agents to give.

For content that needs a human author, the knowledge base gives teams folders and a rich editor — and accepts documents straight from the desktop.

  • Drag-and-drop upload — PDF, Word, CSV, JSON, Markdown and text, up to 20 MB each
  • Rich Markdown editor with tables and code blocks, and a visual ⇄ source toggle
  • Folders by team or topic — the unit of access for agents and workflow steps
  • Usage per item — how often it’s used, and when it was last used
Retrieval quality

Not just “vector search”.

Semantic search alone misses course codes, policy numbers and names. Gecko blends meaning with exact matching, then gives you the tools to tune precision and recall.

Expand

Rewrite the question into broader search variants so relevant content isn’t missed.

Expand User Query

Search

Hybrid by default — vector similarity fused with full-text search using reciprocal rank fusion.

Knowledge · Websites

Rerank

Re-score results with a dedicated reranking model from Cohere, Amazon Bedrock or a compatible endpoint.

Rerank Results

Consolidate

Merge what matters and keep it inside a context budget before it reaches the model.

Consolidate Context

Relevance thresholds

Per-agent minimum scores discard weak results before the model ever sees them.

Semantic chunking

Optionally split website content by meaning rather than fixed length, per site.

Exact-match precision

Full-text search catches the course code, the form number and the building name that embeddings blur.

Transparency

See which content earns its keep.

Every crawled page and knowledge item records how often agents use it. Every answer records the sources behind it. Content ROI stops being a guess.

  • Page tree with embedding status, “Used in N answers” and last-used date
  • Stats tab charting which URLs actually surface in search results
  • One-click exclusions — exclude a page and bulk-remove matching pages already scraped
  • Citations with relevance scores on every answer and every run trace

Point it at your website on Monday; answers cite live pages with relevance scores by Tuesday — and you can see exactly which content nobody needed.

Scoped by design

No agent sees everything by default.

Agents and workflow steps are granted specific knowledge folders and websites. The finance agent reads finance content; the admissions agent never touches HR policy.

Folder-level grants

Choose exactly which knowledge folders each agent or retrieval step can search.

Site-level grants

Select the crawled websites an agent may search, with its own result limit and relevance floor for each.

Role-scoped editing

Custom roles can be limited to specific websites, so a department manages its own content and nothing else.

Start with what you already publish

See your own content answering questions — with sources.

Bring a website or a SharePoint site. We’ll show you how it would ground an agent, and which content would do the heavy lifting.

  • Your content, crawled and cited
  • The scoping and controls it needs
  • A clear view of content gaps
Book a conversation