topics: ai & agents · evaluation & judging
searching ai & agents · evaluation & judging removeMCP servers removeneeds a key remove clear all
- invinoveritas invinoveritas gives an agent a neutral pre-action verdict, a signed post-action proof, and a publicly verifiable on-chain track record for another agent's output. [not the operator's words] 31 skills · needs a key · checked 31m ago 0
- ai-wave AI model changes, benchmark scores and cost estimates. tools/call needs an API key: create one at https://aiwave.elopstudio.com/member and send it as Authorization: Bearer <key>. See https://aiwave.elopstudio.com/llms.txt for when to use which tool. 8 skills · needs a key · checked 23m ago 0
- mumo mumo runs a deliberation across 2-3 frontier models from different labs and returns each model's full response plus a claim map of where they agree and diverge. Reach for it when the cost of being wrong exceeds the cost of deliberating: contested architecture and product decisions, plan and spec rev… 8 skills · needs a key · checked 1h ago 0
- CompletionKit Prompt evals over MCP: run a prompt on your dataset, score each output 1-5 with an LLM judge. [not the operator's words] 54 skills · needs a key · checked 37m ago 0