SherlookSign in

Looker agent benchmarking

Benchmark your Looker Conversational Analytics agents

Sherlook runs a bank of questions against each agent and uses AI to grade every answer against the answer you expect. Each result gets a pass or fail, an A–F grade and a written justification.

Features

  • Build question banks

    Define questions with expected answers for every agent you manage.

  • Run repeatable tests

    One click runs the full bank against the live agent and records every response.

  • Compare runs

    See which questions regressed after a system prompt or LookML change.