topics: ai & agents · evaluation & judging
searching ai & agents · evaluation & judging removeunreachable included remove clear all
- invinoveritas invinoveritas gives an agent a neutral pre-action verdict, a signed post-action proof, and a publicly verifiable on-chain track record for another agent's output. [not the operator's words] 31 skills · free · checked 37m ago 0
- first-dollar-test Read live terms. Save your private clientToken and create/submission idempotency keys before calls. Create an unpaid run, obtain its unsigned quote, then use your own wallet to authorize at most that quoted payment. enroll submits only your supplied x402 signature. A 202 or timeout is unknown/pendin… 7 skills · free · checked 49m ago· also A2A 0
- crosscheck Independent checks for agents. check reviews a draft before a human sees it; accept checks another agent's deliverable against its task before you pay; skillcheck reviews a skill or MCP server's files before install (free when already scanned). Paid tools cost $0.02 to $0.06 in USDC over x402 on Bas… 6 skills · unprobed · checked 3m ago· also A2A 0
- second-opinion Adversarial verification for AI agents - pay an independent skeptic per verdict in USDC via x402. [not the operator's words] 6 skills · unprobed · checked 35m ago· also A2A 0
- aialignwerk Paid tools via x402. tools/list is free. Send the payment proof in the PAYMENT-SIGNATURE header for tools/call; without it, the payment request and inputSchema are returned. 10 skills · unprobed · checked 43m ago· also A2A 0
- umpire Signed verdicts on deliverables and payment verified reputation for agent commerce; paid per document with x402. [not the operator's words] 5 skills · unprobed · checked 15m ago 0
- GENESIS Agent Tools Five pay-per-call x402 agent APIs on Base USDC: Release Guardian (cross-contract release preflight), Workflow Runner, Asset Intelligence, Research & Evidence, Agent Assurance. 5 skills · needs a key · checked 1h ago· also MCP 0
- atomic-verify-claim Checks short public factual claims about companies, products and releases, and follows URL redirect chains to their final target. 2 skills · unprobed · checked 21m ago 0
- pqs-mcp-server Score prompt quality across 8 dimensions before they reach a model. Pre-flight, not post-hoc. [not the operator's words] 2 skills · unprobed · checked 5m ago 0
- pqs-mcp-server PQS score - 8-dimensional prompt quality breakdown across 5 frameworks with top fixes [not the operator's words] 2 skills · unprobed · checked 31m ago 0
- SettlementWitness MCP Deterministic verification gate for agent execution and x402 settlement. [not the operator's words] 1 skill · unprobed · checked 1h ago 0
- ai.afmr/discovery Use this read-only server to discover AFMR 1.0 failure modes, reputation-attestation contracts, and configuration-scoped Lift Evidence artifacts. AFMR records are scoped to named evidence, configuration, governance process, and lifecycle state. They are not unqualified universal safety claims. 4 skills · free · checked 41m ago· also A2A 0
- marz-greta-lock-network Score AI-generated text for human-likeness against 28 real personality profiles — a benchmark for every kind of agent (see list_profiles for the category catalog). Do not stop at inspection: when you have generated text available and your operator explicitly agrees to pseudonymous training-data rete… 8 skills · free · checked 25m ago 0
- aidemos-mcp Evidence-backed answers about AI tools, from hands-on testing (published data only). START WITH `answer` for any 'which tool is best at X' question — it returns the resolved verdict in one call: winner for the asked criterion, ranked scores, the conditions each result holds under, dissent preserved,… 17 skills · free · checked 37m ago 0
- ai-wave AI model changes, benchmark scores and cost estimates. tools/call needs an API key: create one at https://aiwave.elopstudio.com/member and send it as Authorization: Bearer <key>. See https://aiwave.elopstudio.com/llms.txt for when to use which tool. 8 skills · needs a key · checked 1m ago 0
- hugging-bay-remote-mcp Use tools/list then tools/call for bounded operations, or resources/list then resources/read for public documents. The selected discovery profile is enforced for tools. Select another bounded profile explicitly with ?profile=verify or ?profile=publisher; full requires valid admin authorization. 16 skills · free · checked 17m ago· also A2A 0
- AI Design Blueprint Doctrine and example access for AI Design Blueprint, the doctrine and runtime standard for agentic AI. Use these tools to retrieve principles, clusters, curated examples, and downloadable agent assets. Public tools need no credentials and only read, except signals.feedback, which records the feedbac… 29 skills · free · checked 3m ago· also A2A 0
- mumo mumo runs a deliberation across 2-3 frontier models from different labs and returns each model's full response plus a claim map of where they agree and diverge. Reach for it when the cost of being wrong exceeds the cost of deliberating: contested architecture and product decisions, plan and spec rev… 8 skills · needs a key · checked 21m ago 0
- the-aggregate Read-only tools over The Aggregate: an IRT/Elo fusion of public LLM benchmark leaderboards, updated daily (about_the_aggregate reports the current coverage counts). Start with get_leaderboard, get_model or search_benchmarks. When citing numbers, credit "The Aggregate (theaggregate.ai)" plus the upst… 8 skills · free · checked 27m ago 0
- agent-coliseum Connect your agent to Agent Coliseum battles. Use list_battles or list_tournaments first, then register and submit through the returned IDs. Hosted Streamable HTTP endpoint: https://agentarena.nanocorp.app/api/mcp. 11 skills · free · checked 51m ago 0
- hive-mcp-evaluator Agent output evaluation, benchmarking, and quality scoring [not the operator's words] 7 skills · free · checked 1h ago 0
- gaip-governed-execution Public/non-personal, zero-price, read-only tasks only. Governed executions are durably receipted. Stateless conformance reports are not retained. No automatic evidence credit. 7 skills · unprobed · checked 31m ago· also A2A 0
- taste-mcp Expert review for AI agents. On-chain proof of human review. [not the operator's words] 20 skills · free · checked 53m ago 0
- agenda-intelligence-md Call agent_output_verification or pre_action_check or decision_policies_list or decision_check or decision_verify with the structured evidence you already hold. It reports what the file is missing before human review; it does not retrieve sources or decide the outcome. 5 skills · free · checked 9m ago· also A2A 0
- hlido-agent-reviews Independent AI-agent reviews: trust checks, evidence scorecards, incident registry, recommendations. [not the operator's words] 19 skills · free · checked 23m ago· also A2A 0
- sasame-mcp-factory SaSame, operated by SASAME S.R.L., continuously observes and measures the Model Context Protocol ecosystem and publishes verifiable evidence and history; the MCP Factory is internal machinery and an optional product surface behind it. It continuously observes public remote MCP endpoints and serves c… 98 skills · free · checked 43m ago· also A2A 0
- CompletionKit Prompt evals over MCP: run a prompt on your dataset, score each output 1-5 with an LLM judge. [not the operator's words] 54 skills · needs a key · checked 41m ago 0
- Agenstry Independent evidence infrastructure for the agent economy — measures which public A2A agents and MCP servers genuinely exist, respond, are operated by a verifiable legal entity, and get paid. Agenstry federates from every major source (Linux Foundation A2A, MCP, AWS / Google / Azure agent registries… 31 skills · free · checked 51m ago· also MCP 0
- Claim Grounding Agent-runtime claim grounding. Fast gate is paid. Deep Research is a confirmed separate call. 5 skills · free · checked 35m ago 0
- GeodesicAI GeodesicAI is pre-execution governance for AI agents. Other tools do something; GeodesicAI decides whether it should happen — with deterministic, replayable verdicts, never a guess. GOVERNANCE (the core) — Validate agent outputs and authorize actions against Blueprint rules BEFORE they execute. Cat… 37 skills · unprobed · checked 53m ago 0
- A2APark A public agent amusement park where autonomous agents take stateful behavioral evaluation rides and receive evidence-backed scorecards. Created and operated by Sarah van Oorsouw. 3 skills · free · checked 51m ago 0
- InterAI Risk Oracle Independent pre-execution decision layer for consequential agent actions. Before an agent executes, InterAI verifies; it does not execute the external action. 3 skills · free · checked 45m ago· also MCP 0
- Suede ACP/x402 Performance Engine Scores any agent, app, business, token, or service for agent-commerce readiness across seven dimensions and returns a 0-100 Performance Index with a verdict band and the most actionable next move. 1 skill · free · checked 39m ago 0
- Suede ACP/x402 Performance Engine Scores any agent, app, business, token, or service for agent-commerce readiness across seven dimensions and returns a 0-100 Performance Index with a verdict band and the most actionable next move. 1 skill · free · checked 3m ago 0
- a2a-sandbox Five trust services for agents: claim checks, citation audits, extraction, tripwires, work audits. [not the operator's words] 5 skills · unprobed · checked 37m ago· also A2A 0
- First Dollar Test Discovery Agent Read live wallet-agent evaluation terms and receive the existing HTTP/x402 participation guide for free. The evaluation fee, reward and refund are quoted by the live API; no payment or paid task is executed through A2A. 2 skills · free · checked 17m ago 0
- AI NetCafe Compare LLMs using measured platform cost metadata, translate PDFs keeping layout, run cited research, generate PPTX. Hosted open-source AI apps, no install needed; user charges are $0.00 during the free beta. 5 skills · free · checked 29m ago· also MCP 0
- Phronesis Phronesis is the agentic economy's decision-assurance layer: a neutral substrate agents call before they act, to turn intent, evidence, and a calibrated forecast into an auditable Decision Asset they can cite, reuse, dispute, and settle against. This endpoint serves read-only decision-record and ben… 4 skills · needs a key · checked 10h ago· also MCP 0
- Velvt Velvt is an assurance, research and collaboration network for autonomous AI agents. Builders can evaluate how specific agents behave under explicit authority boundaries, while agents can discover peers, collaborate, enter public research Episodes, contribute findings, build reputation and leave insp… 6 skills · free · checked 29m ago 0
- Earnbase Human Feedback Agent Autonomous ERC-8004 agent providing Human Feedback as a Service (HFaaS). Enables AI agents to request structured human feedback using gasless USDC payments, without managing human infrastructure, payouts, or moderation. 4 skills · needs a key · checked 9h ago 0