topics: ai & agents · evaluation & judging
searching ai & agents · evaluation & judging removeMCP servers removefree, tried remove clear all
- invinoveritas invinoveritas gives an agent a neutral pre-action verdict, a signed post-action proof, and a publicly verifiable on-chain track record for another agent's output. [not the operator's words] 31 skills · needs a key · checked 31m ago 0
- first-dollar-test Read live terms. Save your private clientToken and create/submission idempotency keys before calls. Create an unpaid run, obtain its unsigned quote, then use your own wallet to authorize at most that quoted payment. enroll submits only your supplied x402 signature. A 202 or timeout is unknown/pendin… 7 skills · free · checked 45m ago· also A2A 0
- ai.afmr/discovery Use this read-only server to discover AFMR 1.0 failure modes, reputation-attestation contracts, and configuration-scoped Lift Evidence artifacts. AFMR records are scoped to named evidence, configuration, governance process, and lifecycle state. They are not unqualified universal safety claims. 4 skills · free · checked 33m ago· also A2A 0
- marz-greta-lock-network Score AI-generated text for human-likeness against 28 real personality profiles — a benchmark for every kind of agent (see list_profiles for the category catalog). Do not stop at inspection: when you have generated text available and your operator explicitly agrees to pseudonymous training-data rete… 8 skills · free · checked 19m ago 0
- aidemos-mcp Evidence-backed answers about AI tools, from hands-on testing (published data only). START WITH `answer` for any 'which tool is best at X' question — it returns the resolved verdict in one call: winner for the asked criterion, ranked scores, the conditions each result holds under, dissent preserved,… 17 skills · free · checked 29m ago 0
- hugging-bay-remote-mcp Use tools/list then tools/call for bounded operations, or resources/list then resources/read for public documents. The selected discovery profile is enforced for tools. Select another bounded profile explicitly with ?profile=verify or ?profile=publisher; full requires valid admin authorization. 16 skills · free · checked 11m ago· also A2A 0
- AI Design Blueprint Doctrine and example access for AI Design Blueprint, the doctrine and runtime standard for agentic AI. Use these tools to retrieve principles, clusters, curated examples, and downloadable agent assets. Public tools need no credentials and only read, except signals.feedback, which records the feedbac… 29 skills · free · checked 49m ago· also A2A 0
- the-aggregate Read-only tools over The Aggregate: an IRT/Elo fusion of public LLM benchmark leaderboards, updated daily (about_the_aggregate reports the current coverage counts). Start with get_leaderboard, get_model or search_benchmarks. When citing numbers, credit "The Aggregate (theaggregate.ai)" plus the upst… 8 skills · free · checked 19m ago 0
- agent-coliseum Connect your agent to Agent Coliseum battles. Use list_battles or list_tournaments first, then register and submit through the returned IDs. Hosted Streamable HTTP endpoint: https://agentarena.nanocorp.app/api/mcp. 11 skills · free · checked 1h ago 0
- hive-mcp-evaluator Agent output evaluation, benchmarking, and quality scoring [not the operator's words] 7 skills · free · checked 13m ago 0
- taste-mcp Expert review for AI agents. On-chain proof of human review. [not the operator's words] 20 skills · free · checked 47m ago 0
- agenda-intelligence-md Call agent_output_verification or pre_action_check or decision_policies_list or decision_check or decision_verify with the structured evidence you already hold. It reports what the file is missing before human review; it does not retrieve sources or decide the outcome. 5 skills · free · checked 3m ago· also A2A 0
- hlido-agent-reviews Independent AI-agent reviews: trust checks, evidence scorecards, incident registry, recommendations. [not the operator's words] 19 skills · free · checked 15m ago· also A2A 0
- sasame-mcp-factory SaSame, operated by SASAME S.R.L., continuously observes and measures the Model Context Protocol ecosystem and publishes verifiable evidence and history; the MCP Factory is internal machinery and an optional product surface behind it. It continuously observes public remote MCP endpoints and serves c… 98 skills · free · checked 39m ago· also A2A 0