Mast Skills

devops-infra MCP Server

Empirical benchmark comparing agent architectures (single-agent, multi-agent, adaptive) on ProgramDev-v0 and CyberGym tasks. Key finding: adaptive architecture > single-agent > fixed-pipeline multi-agent.

Verified
devops-infradevops-infra
5 views1 stars0 forksMIT

Why This Matters

Discovered via github-topic:mcp and last synced 4mo ago.

Verified
Source
github-topic:mcp
Stars
1
Last synced
4mo ago
Install
Check source

Install

Install instructions not detected yet

Check the source repository for the latest setup steps.

View source instructions
24
Tools
0
Resources
0
Prompts
Standard I/O
Transport

Available Tools (24)

Method

Result

Intervention

ChatDev ProgramDev

Lean

**+31.6pp**

8

*Claude Code Opus and Private Agent MiniMax are 4-rep averages. Others are single rep.* *† 29/30 judged. ‡ 25/30 judged. Skipped tasks counted as non-PASS.* ## CyberGym Results (10 vulnerability tasks) Preliminary results on real-world vulnerability analysis (PoC generation):

Prompt

Hermes GLM

Score

FAILs

Metric

1-rep

1

**Architecture**

0

~10 (file, bash, search)

2

**Model**

Config

PoCs Generated

Model

Provider

Framework

Architecture

Rep

Executability

Rank

Factor

3

**Prompts**

Hermes

Single-agent + tool calling

4

**Middleware**

ChatDev

Fixed pipeline (9 roles)

MiniMax

6/30 (20%)

Baseline

1/4 (25%)

CC

MiniMax

Verbose

**+23.3pp**

r1

30/30 (100%)