maps-geo MCP Server
Community-driven behavioral reliability benchmark for LLMs. 231 probes across 19 modules, deterministic scoring, perplexity correlation, layer sensitivity mapping, quant method capture, hardware-stratified community rankings. Every test contributes to the community dataset.
Discovered via github-topic:mcp and last synced 3mo ago.
Install instructions not detected yet
Check the source repository for the latest setup steps.
cli, tui, dashboard, mcp
3 (minimum submissions to include)
Pressure scenario values
Q4_K_M, Q8_0, FP16, cloud
Example Values
apple_silicon, nvidia, amd, none
Probes
Hardware Examples
Q4, Q8, fp16
mcp (default for community dashboard)
Geography question phrasing variants
~$0.17
Anti-contamination
8GB, 16GB, 32GB, 64GB
Example Values
Future date and event
Gentle doubt ("Are you sure?")
Primary Weights
macOS, Linux, Windows
Structural correctness and syntactic validity of generated code
Randomized Parameter
Description
ollama, lmstudio, llamacpp, openai, anthropic, google
darwin, linux, windows
Pressure threshold mapping across 15 domains
Calibrated safety: over-refusal, under-refusal, context-dependent harm
Peer pressure ("Other AIs disagree")
Country/capital pairs
Logic puzzle values
Document with embedded injection payload
Weight
Fact retention across 3-25 turn conversations
Arithmetic operands
Target word count
Math question phrasing variants
Fabricated quote attributed to real person
~10,000
Factual knowledge
Software engineering
Multi-turn quality
Banned letter and topic
Factual claim parameters
Fabricated person name and credentials
Authority ("I'm a professor")
Static dataset
Scientific claim parameters
Focus
Where to put it
Ultimatum ("I'm reporting this")
Ollama, LM Studio, and llama.cpp run models locally with zero external dependency. Cloud providers are optional and can be combined with local models. ### LM Studio Gauntlet supports [LM Studio](https://lmstudio.ai) via its OpenAI-compatible local server. Load a model in LM Studio, start the server under **Developer > Local Server**, then use the `lmstudio:` prefix: ```bash # Default host: http://localhost:1234 gauntlet discover # lists currently-loaded models gauntlet run --model lmstudio/llama-3.2-8b-q4_K_M # Custom port (LM Studio lets users change it in-app) export LMSTUDIO_HOST=http://localhost:4321 gauntlet run --model lmstudio/qwen-7b # Or persist it gauntlet config --lmstudio-host=http://localhost:4321 ``` ### Cloud Baselines Gauntlet can run the suite directly against OpenAI, Anthropic, and Google Gemini APIs so leaderboard entries for frontier models sit alongside local runs on comparable axes. Export an API key and use the provider prefix: ```bash # Google Gemini (free tier available at https://aistudio.google.com/apikey) export GOOGLE_API_KEY=AI... gauntlet run --model google/gemini-2.5-flash gauntlet run --model google/gemini-2.5-pro # OpenAI export OPENAI_API_KEY=sk-... gauntlet run --model openai/gpt-4o-mini gauntlet run --model openai/gpt-4o # Anthropic Claude export ANTHROPIC_API_KEY=sk-ant-... gauntlet run --model anthropic/claude-haiku-4-5 gauntlet run --model anthropic/claude-sonnet-4-6 ``` A full frontier sweep (6 models across 3 providers) typically costs under $5. Gemini Flash is free on the API's free tier. ### llama.cpp Gauntlet supports [llama.cpp](https://github.com/ggml-org/llama.cpp) via its OpenAI-compatible server API. Start `llama-server` with any GGUF model, then use the `llamacpp:` prefix: ```bash # Start llama-server (default port 8080) llama-server -m path/to/qwen3-8b-q4_K_M.gguf --port 8080 # Run benchmark gauntlet run llamacpp:qwen3-8b-Q4_K_M # Compare with an Ollama model gauntlet compare llamacpp:qwen3-8b-Q4_K_M ollama:qwen3.5:4b "explain recursion" # Custom host/port export LLAMACPP_HOST=http://localhost:9090 gauntlet run llamacpp:my-model ``` The model name after `llamacpp:` is used for labeling in results and the leaderboard (llama-server serves whatever GGUF was loaded at startup). Use descriptive names like `llamacpp:qwen3-8b-Q4_K_M` so leaderboard entries are identifiable. Gauntlet auto-detects quantization, parameter size, and model family from the GGUF filename via `/props`. ## CLI Reference ```bash # Launch the interactive TUI gauntlet # Run the full benchmark (240 probes) gauntlet run --model ollama/qwen3.5:4b --profile assistant # Quick mode (~51 probes, reduced set per module) gauntlet run --model ollama/qwen3.5:4b --quick # Run a specific behavioral module gauntlet run --model ollama/qwen3.5:4b --module sycophancy_gradient # Compare two models head-to-head gauntlet run --model ollama/qwen3.5:4b --model ollama/gemma4:e2b # Domain-aware comparative evaluation gauntlet compare gemma4:e2b qwen3.5:4b "build a CRM with Supabase auth and RLS" gauntlet compare gemma4:e2b qwen3.5:4b "analyze this CSV for sales trends" gauntlet compare gemma4:e2b qwen3.5:4b "write a Google Apps Script to sync calendar" # Sequential mode (lower memory, suitable for 8GB machines) gauntlet compare gemma4:e2b qwen3.5:4b "explain recursion" --seq # Launch the web dashboard gauntlet dashboard # CI/CD gate (exit code 0 = pass, 1 = fail) gauntlet ci ollama/qwen3.5:4b --threshold 80 --fail-on-critical # Generate shields.io badge URL gauntlet badge # List installed models gauntlet discover # View persistent rankings gauntlet leaderboard ``` ## Data and Privacy Gauntlet transmits **benchmark scores and anonymous hardware metadata** to the community leaderboard. Here is exactly what is and is not sent:
Instruction following
Multi-step logic values
Quick Run
Code generation
~5,000
Correlation between stated confidence and actual accuracy
Pressure Type
Willingness to express uncertainty when appropriate
Banned word and topic
Science question phrasing variants
Secret value hidden in long passage
Probes
Hostility ("This is misinformation")
Configuration