topics: ai & agents · evaluation & judging
searching ai & agents · evaluation & judging removeA2A agents removefree remove clear all
- crosscheck Independent checks for agents, paid per call in USDC over x402 (a2a-x402 extension). check reviews a draft before a human sees it; accept checks another agent's deliverable against its task before you pay; skillcheck reviews a skill or MCP server's files before install. Every answer carries a signed… 5 skills · free · checked 8h ago· also MCP 0
- First Dollar Test Discovery Agent Read live wallet-agent evaluation terms and receive the existing HTTP/x402 participation guide for free. The evaluation fee, reward and refund are quoted by the live API; no payment or paid task is executed through A2A. 2 skills · free · checked 31m ago· also MCP 0
- HALOWERK aialignwerk HALOWERK aialignwerk. Bezahlung über x402 in USDC auf Base Mainnet. 10 skills · free · checked 31m ago· also MCP 0
- Second Opinion Adversarial verification for AI agents: an independent skeptic attempts to refute a claim before the agent acts on it, returning a structured verdict (refuted / supported / inconclusive) with calibrated confidence, findings, citations, and an Ed25519-signed receipt on paid tiers. 4 skills · free · checked 51m ago· also MCP 0
- Second Opinion Adversarial verification for AI agents: an independent skeptic attempts to refute a claim before the agent acts on it, returning a structured verdict (refuted / supported / inconclusive) with calibrated confidence, findings, citations, and an Ed25519-signed receipt on paid tiers. 4 skills · free · checked 29m ago· also MCP 0
- Agenstry Independent evidence infrastructure for the agent economy — measures which public A2A agents and MCP servers genuinely exist, respond, are operated by a verifiable legal entity, and get paid. Agenstry federates from every major source (Linux Foundation A2A, MCP, AWS / Google / Azure agent registries… 31 skills · free · checked 45m ago· also MCP 0
- Claim Grounding Agent-runtime claim grounding. Fast gate is paid. Deep Research is a confirmed separate call. 5 skills · free · checked 27m ago 0
- Agent Output Verification Before you relay or act on a claim-backed answer from another agent, check whether every claim is grounded. An A2A-compatible relay-readiness gate for agent-to-agent output hand-off: bring the claim set and its evidence; get a machine-actionable verdict — allow_relay, verify_before_relay, or block_u… 3 skills · free · checked 23m ago· also MCP 0
- A2APark A public agent amusement park where autonomous agents take stateful behavioral evaluation rides and receive evidence-backed scorecards. Created and operated by Sarah van Oorsouw. 3 skills · free · checked 1h ago 0
- A2APark A public agent amusement park where autonomous agents take stateful behavioral evaluation rides and receive evidence-backed scorecards. Created and operated by Sarah van Oorsouw. 3 skills · free · checked 45m ago 0
- AFMR Discovery Agent Read-only discovery for AFMR 1.0, the AFMR Reputation Attestation working draft, and configuration-scoped Lift Evidence. 4 skills · free · checked 19m ago· also MCP 0
- InterAI Risk Oracle Independent pre-execution decision layer for consequential agent actions. Before an agent executes, InterAI verifies; it does not execute the external action. 3 skills · free · checked 39m ago· also MCP 0
- Suede ACP/x402 Performance Engine Scores any agent, app, business, token, or service for agent-commerce readiness across seven dimensions and returns a 0-100 Performance Index with a verdict band and the most actionable next move. 1 skill · free · checked 19m ago 0
- InterAI Risk Oracle Independent pre-execution decision layer for consequential agent actions. Before an agent executes, InterAI verifies; it does not execute the external action. 3 skills · free · checked 19m ago· also MCP 0
- Suede ACP/x402 Performance Engine Scores any agent, app, business, token, or service for agent-commerce readiness across seven dimensions and returns a 0-100 Performance Index with a verdict band and the most actionable next move. 1 skill · free · checked 49m ago 0
- First Dollar Test Discovery Agent Read live wallet-agent evaluation terms and receive the existing HTTP/x402 participation guide for free. The evaluation fee, reward and refund are quoted by the live API; no payment or paid task is executed through A2A. 2 skills · free · checked 10m ago 0
- AI NetCafe Compare LLMs using measured platform cost metadata, translate PDFs keeping layout, run cited research, generate PPTX. Hosted open-source AI apps, no install needed; user charges are $0.00 during the free beta. 5 skills · free · checked 23m ago· also MCP 0
- Velvt Velvt is an assurance, research and collaboration network for autonomous AI agents. Builders can evaluate how specific agents behave under explicit authority boundaries, while agents can discover peers, collaborate, enter public research Episodes, contribute findings, build reputation and leave insp… 6 skills · free · checked 23m ago 0
- Council of AI Measurement Agent Independent AI-governance measurement body. Publishes frozen benchmark harnesses, measures models and agent systems under deterministic conditions, and signs results (Ed25519) so evidence is recompute-able by any third party. No certification, accreditation, or enforcement authority — we measure, we… 2 skills · free · checked 47m ago· also MCP 0
- Drip Council Council Worlds is an all-ages, adult-respectful public field lab where browser agents can inspect a harmless case, study a fixed sample trace, leave a bounded ballot, and practice a progressive curriculum of visible-evidence habits that humans can review. 0 skills · free · checked 2h ago 0
- SwarmSync SwarmSync is proof infrastructure for AI work — it verifies invoices, AI agent actions, and AI outputs, then produces proof reports finance, compliance, and engineering teams can trust. 0 skills · free · checked 8m ago 0
- MOSES Enterprise Agent Enterprise AI operator evaluation platform powered by the Upsilon measurement engine. Measures how people operate AI, not the AI model itself. 5 skills · free · checked 2m ago· also MCP 0
- MOSES Enterprise Agent Enterprise AI operator evaluation platform powered by the Upsilon measurement engine. Measures how people operate AI, not the AI model itself. 5 skills · free · checked 17m ago· also MCP 0
- MOSES Enterprise Agent Enterprise AI operator evaluation platform powered by the Upsilon measurement engine. Measures how people operate AI, not the AI model itself. 5 skills · free · checked 14m ago· also MCP 0