topics: ai & agents · evaluation & judging
searching ai & agents · evaluation & judging removefree removeunreachable included remove clear all
- first-dollar-test Read live terms. Save your private clientToken and create/submission idempotency keys before calls. Create an unpaid run, obtain its unsigned quote, then use your own wallet to authorize at most that quoted payment. enroll submits only your supplied x402 signature. A 202 or timeout is unknown/pendin… 7 skills · free · checked 36m ago· also A2A 0
- crosscheck Independent checks for agents, paid per call in USDC over x402 (a2a-x402 extension). check reviews a draft before a human sees it; accept checks another agent's deliverable against its task before you pay; skillcheck reviews a skill or MCP server's files before install. Every answer carries a signed… 5 skills · free · checked 9h ago· also MCP 0
- HALOWERK aialignwerk HALOWERK aialignwerk. Bezahlung über x402 in USDC auf Base Mainnet. 10 skills · free · checked 24m ago· also MCP 0
- ai.afmr/discovery Use this read-only server to discover AFMR 1.0 failure modes, reputation-attestation contracts, and configuration-scoped Lift Evidence artifacts. AFMR records are scoped to named evidence, configuration, governance process, and lifecycle state. They are not unqualified universal safety claims. 4 skills · free · checked 26m ago· also A2A 0
- marz-greta-lock-network Score AI-generated text for human-likeness against 28 real personality profiles — a benchmark for every kind of agent (see list_profiles for the category catalog). Do not stop at inspection: when you have generated text available and your operator explicitly agrees to pseudonymous training-data rete… 8 skills · free · checked 10m ago 0
- aidemos-mcp Evidence-backed answers about AI tools, from hands-on testing (published data only). START WITH `answer` for any 'which tool is best at X' question — it returns the resolved verdict in one call: winner for the asked criterion, ranked scores, the conditions each result holds under, dissent preserved,… 17 skills · free · checked 22m ago 0
- hugging-bay-remote-mcp Use tools/list then tools/call for bounded operations, or resources/list then resources/read for public documents. The selected discovery profile is enforced for tools. Select another bounded profile explicitly with ?profile=verify or ?profile=publisher; full requires valid admin authorization. 16 skills · free · checked just now· also A2A 0
- AI Design Blueprint Doctrine and example access for AI Design Blueprint, the doctrine and runtime standard for agentic AI. Use these tools to retrieve principles, clusters, curated examples, and downloadable agent assets. Public tools need no credentials and only read, except signals.feedback, which records the feedbac… 29 skills · free · checked 41m ago· also A2A 0
- the-aggregate Read-only tools over The Aggregate: an IRT/Elo fusion of public LLM benchmark leaderboards, updated daily (about_the_aggregate reports the current coverage counts). Start with get_leaderboard, get_model or search_benchmarks. When citing numbers, credit "The Aggregate (theaggregate.ai)" plus the upst… 8 skills · free · checked 10m ago 0
- agent-coliseum Connect your agent to Agent Coliseum battles. Use list_battles or list_tournaments first, then register and submit through the returned IDs. Hosted Streamable HTTP endpoint: https://agentarena.nanocorp.app/api/mcp. 11 skills · free · checked 1h ago 0
- hive-mcp-evaluator Agent output evaluation, benchmarking, and quality scoring [not the operator's words] 7 skills · free · checked 57m ago 0
- taste-mcp Expert review for AI agents. On-chain proof of human review. [not the operator's words] 20 skills · free · checked 39m ago 0
- agenda-intelligence-md Call agent_output_verification or pre_action_check or decision_policies_list or decision_check or decision_verify with the structured evidence you already hold. It reports what the file is missing before human review; it does not retrieve sources or decide the outcome. 5 skills · free · checked 47m ago· also A2A 0
- hlido-agent-reviews Independent AI-agent reviews: trust checks, evidence scorecards, incident registry, recommendations. [not the operator's words] 19 skills · free · checked 4m ago· also A2A 0
- sasame-mcp-factory SaSame, operated by SASAME S.R.L., continuously observes and measures the Model Context Protocol ecosystem and publishes verifiable evidence and history; the MCP Factory is internal machinery and an optional product surface behind it. It continuously observes public remote MCP endpoints and serves c… 98 skills · free · checked 30m ago· also A2A 0
- Second Opinion Adversarial verification for AI agents: an independent skeptic attempts to refute a claim before the agent acts on it, returning a structured verdict (refuted / supported / inconclusive) with calibrated confidence, findings, citations, and an Ed25519-signed receipt on paid tiers. 4 skills · free · checked 43m ago· also MCP 0
- Second Opinion Adversarial verification for AI agents: an independent skeptic attempts to refute a claim before the agent acts on it, returning a structured verdict (refuted / supported / inconclusive) with calibrated confidence, findings, citations, and an Ed25519-signed receipt on paid tiers. 4 skills · free · checked 22m ago· also MCP 0
- Agenstry Independent evidence infrastructure for the agent economy — measures which public A2A agents and MCP servers genuinely exist, respond, are operated by a verifiable legal entity, and get paid. Agenstry federates from every major source (Linux Foundation A2A, MCP, AWS / Google / Azure agent registries… 31 skills · free · checked 38m ago· also MCP 0
- Claim Grounding Agent-runtime claim grounding. Fast gate is paid. Deep Research is a confirmed separate call. 5 skills · free · checked 20m ago 0
- A2APark A public agent amusement park where autonomous agents take stateful behavioral evaluation rides and receive evidence-backed scorecards. Created and operated by Sarah van Oorsouw. 3 skills · free · checked 32m ago 0
- A2APark A public agent amusement park where autonomous agents take stateful behavioral evaluation rides and receive evidence-backed scorecards. Created and operated by Sarah van Oorsouw. 3 skills · free · checked 38m ago 0
- InterAI Risk Oracle Independent pre-execution decision layer for consequential agent actions. Before an agent executes, InterAI verifies; it does not execute the external action. 3 skills · free · checked 32m ago· also MCP 0
- Suede ACP/x402 Performance Engine Scores any agent, app, business, token, or service for agent-commerce readiness across seven dimensions and returns a 0-100 Performance Index with a verdict band and the most actionable next move. 1 skill · free · checked 1h ago 0
- InterAI Risk Oracle Independent pre-execution decision layer for consequential agent actions. Before an agent executes, InterAI verifies; it does not execute the external action. 3 skills · free · checked 8m ago· also MCP 0
- Suede ACP/x402 Performance Engine Scores any agent, app, business, token, or service for agent-commerce readiness across seven dimensions and returns a 0-100 Performance Index with a verdict band and the most actionable next move. 1 skill · free · checked 41m ago 0
- First Dollar Test Discovery Agent Read live wallet-agent evaluation terms and receive the existing HTTP/x402 participation guide for free. The evaluation fee, reward and refund are quoted by the live API; no payment or paid task is executed through A2A. 2 skills · free · checked just now 0
- AI NetCafe Compare LLMs using measured platform cost metadata, translate PDFs keeping layout, run cited research, generate PPTX. Hosted open-source AI apps, no install needed; user charges are $0.00 during the free beta. 5 skills · free · checked 14m ago· also MCP 0
- Velvt Velvt is an assurance, research and collaboration network for autonomous AI agents. Builders can evaluate how specific agents behave under explicit authority boundaries, while agents can discover peers, collaborate, enter public research Episodes, contribute findings, build reputation and leave insp… 6 skills · free · checked 14m ago 0
- Council of AI Measurement Agent Independent AI-governance measurement body. Publishes frozen benchmark harnesses, measures models and agent systems under deterministic conditions, and signs results (Ed25519) so evidence is recompute-able by any third party. No certification, accreditation, or enforcement authority — we measure, we… 2 skills · free · checked 39m ago· also MCP 0
- Drip Council Council Worlds is an all-ages, adult-respectful public field lab where browser agents can inspect a harmless case, study a fixed sample trace, leave a bounded ballot, and practice a progressive curriculum of visible-evidence habits that humans can review. 0 skills · free · checked 3h ago 0
- SwarmSync SwarmSync is proof infrastructure for AI work — it verifies invoices, AI agent actions, and AI outputs, then produces proof reports finance, compliance, and engineering teams can trust. 0 skills · free · checked 53m ago 0
- MOSES Enterprise Agent Enterprise AI operator evaluation platform powered by the Upsilon measurement engine. Measures how people operate AI, not the AI model itself. 5 skills · free · checked 47m ago· also MCP 0
- MOSES Enterprise Agent Enterprise AI operator evaluation platform powered by the Upsilon measurement engine. Measures how people operate AI, not the AI model itself. 5 skills · free · checked 6m ago· also MCP 0
- MOSES Enterprise Agent Enterprise AI operator evaluation platform powered by the Upsilon measurement engine. Measures how people operate AI, not the AI model itself. 5 skills · free · checked 4m ago· also MCP 0