Gauntlet

maps-geo MCP Server

Community-driven behavioral reliability benchmark for LLMs. 231 probes across 19 modules, deterministic scoring, perplexity correlation, layer sensitivity mapping, quant method capture, hardware-stratified community rankings. Every test contributes to the community dataset.

Verified
maps-geomaps-geo
4 views6 stars0 forksMIT

Why This Matters

Discovered via github-topic:mcp and last synced 3mo ago.

Verified
Source
github-topic:mcp
Stars
6
Last synced
3mo ago
Install
Check source

Install

Install instructions not detected yet

Check the source repository for the latest setup steps.

View source instructions
64
Tools
0
Resources
0
Prompts
Standard I/O
Transport

Available Tools (64)

source

cli, tui, dashboard, mcp

min_tests

3 (minimum submissions to include)

_FACTORY_SYCOPHANCY

Pressure scenario values

Quantization

Q4_K_M, Q8_0, FP16, cloud

Collected

Example Values

gpu_class

apple_silicon, nvidia, amd, none

Category

Probes

Tier

Hardware Examples

quantization

Q4, Q8, fp16

exclude_source

mcp (default for community dashboard)

_FACTORY_CONSISTENCY_GEOGRAPHY

Geography question phrasing variants

GPT-4o

~$0.17

Multi-turn

Anti-contamination

RAM

8GB, 16GB, 32GB, 64GB

Parameter

Example Values

_FACTORY_FUTURE_PREDICTION

Future date and event

1

Gentle doubt ("Are you sure?")

Profile

Primary Weights

OS

macOS, Linux, Windows

6

Structural correctness and syntactic validity of generated code

Factory

Randomized Parameter

Endpoint

Description

provider

ollama, lmstudio, llamacpp, openai, anthropic, google

os_platform

darwin, linux, windows

15

Pressure threshold mapping across 15 domains

18

Calibrated safety: over-refusal, under-refusal, context-dependent harm

2

Peer pressure ("Other AIs disagree")

_FACTORY_CAPITALS

Country/capital pairs

_FACTORY_SYCOPHANCY_LOGIC

Logic puzzle values

_FACTORY_INJECTION_DOC

Document with embedded injection payload

Severity

Weight

8

Fact retention across 3-25 turn conversations

_FACTORY_MATH

Arithmetic operands

_FACTORY_WORD_COUNT

Target word count

_FACTORY_CONSISTENCY_MATH

Math question phrasing variants

_FACTORY_FAKE_QUOTE

Fabricated quote attributed to real person

167

~10,000

MMLU

Factual knowledge

SWE-bench

Software engineering

MT-Bench

Multi-turn quality

_FACTORY_FORBIDDEN_LETTER

Banned letter and topic

_FACTORY_SYCOPHANCY_FACTUAL

Factual claim parameters

_FACTORY_FAKE_ENTITY

Fabricated person name and credentials

3

Authority ("I'm a professor")

No

Static dataset

_FACTORY_SYCOPHANCY_SCIENCE

Scientific claim parameters

Framework

Focus

Method

Where to put it

5

Ultimatum ("I'm reporting this")

Pay-per-use

Ollama, LM Studio, and llama.cpp run models locally with zero external dependency. Cloud providers are optional and can be combined with local models. ### LM Studio Gauntlet supports [LM Studio](https://lmstudio.ai) via its OpenAI-compatible local server. Load a model in LM Studio, start the server under **Developer > Local Server**, then use the `lmstudio:` prefix: ```bash # Default host: http://localhost:1234 gauntlet discover # lists currently-loaded models gauntlet run --model lmstudio/llama-3.2-8b-q4_K_M # Custom port (LM Studio lets users change it in-app) export LMSTUDIO_HOST=http://localhost:4321 gauntlet run --model lmstudio/qwen-7b # Or persist it gauntlet config --lmstudio-host=http://localhost:4321 ``` ### Cloud Baselines Gauntlet can run the suite directly against OpenAI, Anthropic, and Google Gemini APIs so leaderboard entries for frontier models sit alongside local runs on comparable axes. Export an API key and use the provider prefix: ```bash # Google Gemini (free tier available at https://aistudio.google.com/apikey) export GOOGLE_API_KEY=AI... gauntlet run --model google/gemini-2.5-flash gauntlet run --model google/gemini-2.5-pro # OpenAI export OPENAI_API_KEY=sk-... gauntlet run --model openai/gpt-4o-mini gauntlet run --model openai/gpt-4o # Anthropic Claude export ANTHROPIC_API_KEY=sk-ant-... gauntlet run --model anthropic/claude-haiku-4-5 gauntlet run --model anthropic/claude-sonnet-4-6 ``` A full frontier sweep (6 models across 3 providers) typically costs under $5. Gemini Flash is free on the API's free tier. ### llama.cpp Gauntlet supports [llama.cpp](https://github.com/ggml-org/llama.cpp) via its OpenAI-compatible server API. Start `llama-server` with any GGUF model, then use the `llamacpp:` prefix: ```bash # Start llama-server (default port 8080) llama-server -m path/to/qwen3-8b-q4_K_M.gguf --port 8080 # Run benchmark gauntlet run llamacpp:qwen3-8b-Q4_K_M # Compare with an Ollama model gauntlet compare llamacpp:qwen3-8b-Q4_K_M ollama:qwen3.5:4b "explain recursion" # Custom host/port export LLAMACPP_HOST=http://localhost:9090 gauntlet run llamacpp:my-model ``` The model name after `llamacpp:` is used for labeling in results and the leaderboard (llama-server serves whatever GGUF was loaded at startup). Use descriptive names like `llamacpp:qwen3-8b-Q4_K_M` so leaderboard entries are identifiable. Gauntlet auto-detects quantization, parameter size, and model family from the GGUF filename via `/props`. ## CLI Reference ```bash # Launch the interactive TUI gauntlet # Run the full benchmark (240 probes) gauntlet run --model ollama/qwen3.5:4b --profile assistant # Quick mode (~51 probes, reduced set per module) gauntlet run --model ollama/qwen3.5:4b --quick # Run a specific behavioral module gauntlet run --model ollama/qwen3.5:4b --module sycophancy_gradient # Compare two models head-to-head gauntlet run --model ollama/qwen3.5:4b --model ollama/gemma4:e2b # Domain-aware comparative evaluation gauntlet compare gemma4:e2b qwen3.5:4b "build a CRM with Supabase auth and RLS" gauntlet compare gemma4:e2b qwen3.5:4b "analyze this CSV for sales trends" gauntlet compare gemma4:e2b qwen3.5:4b "write a Google Apps Script to sync calendar" # Sequential mode (lower memory, suitable for 8GB machines) gauntlet compare gemma4:e2b qwen3.5:4b "explain recursion" --seq # Launch the web dashboard gauntlet dashboard # CI/CD gate (exit code 0 = pass, 1 = fail) gauntlet ci ollama/qwen3.5:4b --threshold 80 --fail-on-critical # Generate shields.io badge URL gauntlet badge # List installed models gauntlet discover # View persistent rankings gauntlet leaderboard ``` ## Data and Privacy Gauntlet transmits **benchmark scores and anonymous hardware metadata** to the community leaderboard. Here is exactly what is and is not sent:

AlpacaEval

Instruction following

_FACTORY_LOGIC_CHAIN

Multi-step logic values

Model

Quick Run

HumanEval

Code generation

78

~5,000

10

Correlation between stated confidence and actual accuracy

Level

Pressure Type

12

Willingness to express uncertainty when appropriate

_FACTORY_FORBIDDEN_WORD

Banned word and topic

_FACTORY_CONSISTENCY_SCIENCE

Science question phrasing variants

_FACTORY_NEEDLE

Secret value hidden in long passage

Suite

Probes

4

Hostility ("This is misinformation")

Provider

Configuration