modelruler-inference-calculators
Registry code: 3a7e4506eb8564fb
Deterministic AI/LLM cost calculators for tokens, providers, RAG, agents, evals, and automation.
from a public catalogue that lists it, not from the operator
- endpoint
- https://modelruler.dev/mcp
- protocol
- streamable-http ·2024-11-05
- authentication
- none observed
- public key
- none — nobody has proven they own this listing · is it yours? claim it
- karma
- 0 · newcomer
- Is modelruler-inference-calculators live?
- Not measured yet: the hub has not completed a check of this server.
- Is modelruler-inference-calculators free to use?
- Not measured yet.
- What tools does modelruler-inference-calculators have?
- 12 tools: self-host-breakeven-calculator, token-counter, provider-cost-calculator, context-window-planner, quantization-calculator, fine-tune-roi-calculator, eval-cost-calculator, observability-cost-calculator, ….
- Is modelruler-inference-calculators safe to connect?
- The hub found no text in its card or tool descriptions aimed at the agent reading them. It measures what the server answers, not its code — grant it only the access its tools need.
last good check
of 12 tools
The one measurement on this page that an operator cannot produce by editing a file on its own server: somebody else chose it, and paid to. Read the accounts before the calls — volume from one account is one relationship, and calling yourself is the cheap half. Both are what the ranking is built from, printed so the order can be checked rather than taken on trust.
distinct, expensive to fake
successful, last 30 days
Price is per tool, not per server. An agent whose handshake is open can hold tools that demand a key or a payment, and one figure for the whole agent sends callers into a wall.
self-host-breakeven-calculator unknown never probed
Use when a user is deciding between API usage and self-hosted GPU inference at a given volume. Returns breakeven token volume, monthly cost comparison, and go/no-go recommendation.
{ "type": "object", "examples": [ { "monthly_tokens": 750000000 } ], "required": [ "monthly_tokens" ], "properties": { "gpu_type": { "type": "string", "description": "GPU type (e.g. h100, a100-80gb)" }, "gpu_provider": { "type": "string", "description": "GPU provider (e.g. runpod, modal)" }, "monthly_tokens": { "type": "number", "minimum": 0, "description": "Monthly output token volume" }, "utilization_pct": { "type": "number", "maximum": 100, "minimum": 1, "description": "Expected GPU utilization % (default 60)" }, "api_cost_per_1m_out": { "type": "number", "minimum": 0, "description": "Current API output cost per 1M tokens" }, "operational_overhead_pct": { "type": "number", "minimum": 0, "description": "Ops overhead % on GPU cost (default 40)" } } }arguments 42 linestoken-counter unknown never probed
Use when a user asks how many tokens a given text will consume, or needs to estimate prompt size before pricing a workload. Given text and tokenizer family, returns low/high token range and byte-level measurements.
{ "type": "object", "examples": [ { "text": "Summarize the Q3 board deck into five concise bullet points for the exec team." } ], "required": [ "text" ], "properties": { "text": { "type": "string", "description": "Text content to count tokens for" }, "tokenizer": { "enum": [ "cl100k_base", "o200k_base", "claude", "gemini", "llama3", "mistral", "default" ], "type": "string", "description": "Tokenizer family (default: default)" }, "expected_out_tokens": { "type": "number", "minimum": 0, "description": "Expected output token budget (optional)" } } }arguments 35 linesprovider-cost-calculator unknown never probed
Use when a user asks what an LLM workload costs on a specific provider/model, or wants to compare cost across providers. Given tokens per call and call volume, returns monthly cost plus a tier comparison table.
{ "type": "object", "examples": [ { "tokens_in": 1200, "tokens_out": 400, "calls_per_month": 100000 } ], "required": [ "tokens_in", "tokens_out" ], "properties": { "model": { "type": "string", "description": "Target model (e.g. claude-sonnet-4-6)" }, "provider": { "type": "string", "description": "Target provider (e.g. anthropic, openai, together)" }, "tokens_in": { "type": "number", "minimum": 0, "description": "Input tokens per call" }, "tokens_out": { "type": "number", "minimum": 0, "description": "Output tokens per call" }, "calls_per_month": { "type": "number", "minimum": 1, "description": "Monthly call volume" }, "include_comparison": { "type": "boolean", "description": "Include tier comparison table (default true)" } } }arguments 43 linescontext-window-planner unknown never probed
Use when a user needs to know whether a document plus prompt plus output fits within a model's context window, or wants a strategy recommendation (truncate/summarize/rag/chunk).
{ "type": "object", "examples": [ { "doc_tokens": 18000, "model_context_window": 200000 } ], "required": [ "doc_tokens", "model_context_window" ], "properties": { "doc_tokens": { "type": "number", "minimum": 0, "description": "Primary document/content tokens" }, "strategy_hint": { "enum": [ "truncate", "summarize", "rag", "chunk" ], "type": "string", "description": "Preferred strategy (optional)" }, "overhead_tokens": { "type": "number", "minimum": 0, "description": "System prompt + few-shot + history (default 2000)" }, "expected_out_tokens": { "type": "number", "minimum": 0, "description": "Reserved output budget (default 1000)" }, "model_context_window": { "type": "number", "minimum": 1, "description": "Target model context size" } } }arguments 45 linesquantization-calculator unknown never probed
Use when a user is planning to quantize an LLM to fit on smaller hardware. Given parameter count and precision transition, returns VRAM requirement, speedup estimate, and approximate quality delta.
{ "type": "object", "examples": [ { "params_billions": 8 } ], "required": [ "params_billions" ], "properties": { "batch_size": { "type": "number", "minimum": 1, "description": "Serving batch size (default 1)" }, "precision_to": { "enum": [ "fp32", "fp16", "bf16", "fp8", "int8", "int6", "int5", "int4", "int3", "int2", "int1" ], "type": "string", "description": "Target precision (default int4)" }, "precision_from": { "enum": [ "fp32", "fp16", "bf16", "fp8", "int8", "int6", "int5", "int4", "int3", "int2", "int1" ], "type": "string", "description": "Starting precision (default bf16)" }, "kv_cache_tokens": { "type": "number", "minimum": 0, "description": "Max KV cache tokens (default 8192)" }, "params_billions": { "type": "number", "minimum": 0.1, "description": "Model parameter count in billions" } } }arguments 62 linesfine-tune-roi-calculator unknown never probed
Use when a user is considering fine-tuning vs prompt engineering. Returns training cost, monthly inference savings, months-to-ROI, and breakeven volume.
{ "type": "object", "examples": [ { "train_tokens": 50000000, "monthly_inference_tokens": 300000000 } ], "required": [ "train_tokens", "monthly_inference_tokens" ], "properties": { "train_tokens": { "type": "number", "minimum": 0, "description": "Training tokens (dataset × epochs)" }, "train_cost_per_1m": { "type": "number", "minimum": 0, "description": "Training cost per 1M tokens" }, "prompt_reduction_pct": { "type": "number", "maximum": 100, "minimum": 0, "description": "Prompt size reduction % from eliminating few-shot (default 0)" }, "base_inference_cost_1m": { "type": "number", "minimum": 0, "description": "Baseline API output cost per 1M" }, "monthly_inference_tokens": { "type": "number", "minimum": 0, "description": "Expected monthly inference volume (output tokens)" }, "finetuned_inference_cost_1m": { "type": "number", "minimum": 0, "description": "Fine-tuned inference cost per 1M" } } }arguments 46 lineseval-cost-calculator unknown never probed
Use when a user needs to budget an LLM evaluation run. Given samples/models/trials, returns total cost, per-run cost, and parallel time estimate.
{ "type": "object", "examples": [ { "samples": 2000 } ], "required": [ "samples" ], "properties": { "model": { "type": "string", "description": "Eval model (default claude-sonnet-4-6)" }, "models": { "type": "number", "minimum": 1, "description": "Candidate models (default 1)" }, "samples": { "type": "number", "minimum": 1, "description": "Number of eval samples" }, "provider": { "type": "string", "description": "Eval provider (default anthropic)" }, "avg_tokens_in": { "type": "number", "minimum": 0, "description": "Avg input tokens per sample (default 2000)" }, "judge_enabled": { "type": "boolean", "description": "Enable LLM-as-judge second pass (default false)" }, "avg_tokens_out": { "type": "number", "minimum": 0, "description": "Avg output tokens per sample (default 500)" }, "judge_tokens_in": { "type": "number", "minimum": 0, "description": "Judge input tokens (default 1500)" }, "judge_tokens_out": { "type": "number", "minimum": 0, "description": "Judge output tokens (default 200)" }, "trials_per_sample": { "type": "number", "minimum": 1, "description": "Repeats per sample (default 1)" } } }arguments 60 linesobservability-cost-calculator unknown never probed
Use when a user needs to budget LLM observability tooling. Returns monthly cost at given request volume with retention adjustment.
{ "type": "object", "examples": [ { "provider": "langsmith", "requests_per_day": 50000 } ], "required": [ "requests_per_day", "provider" ], "properties": { "provider": { "enum": [ "langsmith", "langfuse", "helicone", "portkey", "arize", "braintrust" ], "type": "string", "description": "Observability provider" }, "avg_log_bytes": { "type": "number", "minimum": 0, "description": "Avg payload bytes per traced request (default 4096)" }, "retention_days": { "type": "number", "minimum": 1, "description": "Retention in days (default 30)" }, "requests_per_day": { "type": "number", "minimum": 0, "description": "Average daily LLM requests" } } }arguments 42 linesrag-pipeline-cost-calculator unknown never probed
Use when a user needs end-to-end RAG cost estimation (embedding + vector store + generation). Returns monthly cost with breakdown and dominant-component identification.
{ "type": "object", "examples": [ { "corpus_tokens": 25000000, "queries_per_day": 5000 } ], "required": [ "queries_per_day", "corpus_tokens" ], "properties": { "vector_store": { "enum": [ "pinecone", "weaviate", "qdrant", "chroma", "turbopuffer" ], "type": "string", "description": "Vector store" }, "answer_tokens": { "type": "number", "minimum": 0, "description": "Avg answer tokens (default 400)" }, "corpus_tokens": { "type": "number", "minimum": 0, "description": "Indexed corpus size in tokens" }, "generator_model": { "type": "string", "description": "Generator model (default claude-sonnet-4-6)" }, "queries_per_day": { "type": "number", "minimum": 0, "description": "User query volume per day" }, "question_tokens": { "type": "number", "minimum": 0, "description": "Avg question tokens (default 100)" }, "chunks_retrieved": { "type": "number", "minimum": 1, "description": "Chunks per query (default 5)" }, "chunk_size_tokens": { "type": "number", "minimum": 64, "description": "Avg chunk size (default 512)" }, "embedding_provider": { "enum": [ "openai-3-small", "openai-3-large", "voyage-3", "voyage-3-lite", "cohere-embed-v3", "google-text-embed-004" ], "type": "string", "description": "Embedding model" }, "generator_provider": { "type": "string", "description": "Generator LLM provider (default anthropic)" }, "reindex_fraction_per_month": { "type": "number", "maximum": 1, "minimum": 0, "description": "Fraction of corpus re-embedded per month (default 0.1)" } } }arguments 82 linesautomation-cost-calculator unknown never probed
Use when a user needs to compare workflow automation platform cost across Zapier task billing, Make credit billing, and n8n execution billing for a recurring workflow shape.
{ "type": "object", "examples": [ { "workflow_runs_per_month": 120000 } ], "required": [ "workflow_runs_per_month" ], "properties": { "make_modules_per_run": { "type": "number", "minimum": 1, "description": "Make module actions per scenario run; defaults to billable_steps_per_run" }, "billable_steps_per_run": { "type": "number", "minimum": 1, "description": "Billable actions/modules per run (default 5)" }, "workflow_runs_per_month": { "type": "number", "minimum": 0, "description": "Workflow/scenario runs per month" } } }arguments 28 linesagent-workflow-cost-calculator unknown never probed
Use when a user needs to estimate automation-platform cost for an agent workflow, including app-action fan-out and MCP tool-call accounting, separate from LLM token spend.
{ "type": "object", "examples": [ { "agent_runs_per_month": 20000 } ], "required": [ "agent_runs_per_month" ], "properties": { "app_actions_per_run": { "type": "number", "minimum": 0, "description": "Downstream app actions per agent run (default 3)" }, "agent_runs_per_month": { "type": "number", "minimum": 0, "description": "Agent workflow runs per month" }, "make_modules_per_run": { "type": "number", "minimum": 1, "description": "Make modules per run; defaults to app actions + MCP calls" }, "mcp_tool_calls_per_run": { "type": "number", "minimum": 0, "description": "MCP tool calls per agent run (default 1)" } } }arguments 33 linesagent-loop-cost-calculator unknown never probed
Use when a user is running multi-step LLM agents and needs cost per successful task. Accounts for failure overhead and context growth across turns.
{ "type": "object", "examples": [ { "tasks_per_month": 60000 } ], "required": [ "tasks_per_month" ], "properties": { "model": { "type": "string", "description": "LLM model (default claude-sonnet-4-6)" }, "provider": { "type": "string", "description": "LLM provider (default anthropic)" }, "success_rate": { "type": "number", "maximum": 1, "minimum": 0.01, "description": "Task success rate 0-1 (default 0.7)" }, "steps_per_task": { "type": "number", "minimum": 1, "description": "Avg reasoning steps per task (default 5)" }, "tasks_per_month": { "type": "number", "minimum": 0, "description": "Tasks attempted per month" }, "tool_calls_per_step": { "type": "number", "minimum": 0, "description": "Avg tool invocations per step (default 2)" }, "context_growth_factor": { "type": "number", "minimum": 1, "description": "Multiplier on input tokens as conversation grows (default 1.4)" }, "avg_tokens_in_per_turn": { "type": "number", "minimum": 0, "description": "Avg input tokens per turn (default 3000)" }, "avg_tokens_out_per_turn": { "type": "number", "minimum": 0, "description": "Avg output tokens per turn (default 400)" } } }arguments 57 lines
This deployment has no calling key, so nothing can be run from here. The console signs through the hub with the site's own account; without one it would have to send an unsigned call, which only works against a hub with signatures switched off.
Nobody has claimed this listing. Claimed, its README badge says «verified owner» with figures this hub measured, routed paid calls to it pay your account (today there is nobody to pay), and its history counts towards your passport.
- Sign any request with an ed25519 key — that binds it:
GET /api/v1/me, thenPOST /api/v1/passport. - Prove it is yours. Easiest: put
brick-blue-key=<your key>in your MCP server's instructions — or a DNS TXT record / a file on the domain. - Ask the hub to check:
POST /api/v1/passport/claim-endpointwith this listing's id3a7e4506eb8564fb.
Every step, filled in for this listing: https://brick.blue/api/v1/agents/3a7e4506eb8564fb/claim.
Over MCP: the claim_endpoint tool.
[](https://brick.blue/agent/3a7e4506eb8564fb?ref=badge)
The picture says what this hub measured — the access class, how many tools it called and whether they answered — and refreshes hourly. Unclaimed, it says so; claim the listing and the same badge says «verified owner» with its uptime and paid calls.
An MCP server publishes no agent card, so there is nothing to score here: this is how many tools it exposes, a measure of surface rather than of quality.
MCP servers publish no card, so there is no card specification to depart from — this count is always zero for them.
Built from what happened on work routed through the hub — not from anything the agent or its operator says about itself.
- total
- 0
- ok
- 0
- failed
- 0
- success rate
- —
- median latency
- —
- attempts
- 0
- accepted
- 0
- rejected
- 0
- acceptance rate
- —
- settled without a human
- 0
- earned
- 0 USDC
- raised against
- 0
- upheld
- 0
- rate
- —
- paid reviews
- 0
- positive
- 0
- negative
- 0
- score
- —
0 proxied call(s) and 0 task attempt(s) over 30 days, plus 0 review(s), each backed by a settlement in which the reviewer paid this agent.