pytest-skill-engineering

Test-Driven Skill Engineering for GitHub Copilot

Test MCP servers, CLI tools, Agent Skills, and custom agents using the real GitHub Copilot coding agent. Write tests as prompts, run them against actual Copilot sessions, and get AI-powered insights on what to fix.

Why?

Your MCP server passes all unit tests. Then a user tries it in GitHub Copilot and:

Copilot picks the wrong tool
Passes garbage parameters
Can't recover from errors
Ignores your skill's instructions

Why? Because you tested the code, not the AI interface.

For LLMs, your API isn't functions and types — it's tool descriptions, Agent Skills, custom agent instructions, and schemas. These are what GitHub Copilot actually sees. Traditional tests can't validate them.

The key insight: your test is a prompt. You write what a user would say ("What's my checking balance?"), and Copilot figures out how to use your tools. If it can't, your AI interface needs work.

What This Tests

pytest-skill-engineering validates the full skill engineering stack that ships with your MCP server:

MCP Server Tools — Can Copilot discover and call your tools correctly?
Agent Skills (agentskills.io spec-compliant) — Does domain knowledge improve performance?
Custom Agents (.agent.md files) — Do your specialist instructions trigger proper subagent dispatch?
MCP Prompt Templates — Do server-side templates produce the right behavior?
CLI Tools — Can Copilot use command-line interfaces effectively?

Plus A/B testing, multi-turn sessions, clarification detection, and AI-powered reports that tell you exactly what to fix.

How It Works

Write tests as prompts. Run them with the real GitHub Copilot coding agent. Assert on what happened:

from pytest_skill_engineering.copilot import CopilotEval

async def test_balance_query(copilot_eval):
    agent = CopilotEval(
        skill_directories=["skills/banking-advisor"],
        max_turns=10,
    )
    result = await copilot_eval(agent, "What's my checking balance?")
    
    assert result.success
    assert result.tool_was_called("get_balance")

The workflow:

Write a test — a prompt that describes what a user would say
Run it — GitHub Copilot tries to use your tools
Fix the interface — improve tool descriptions, skills, or agent instructions until it passes
AI analysis tells you what to optimize — cost, redundant calls, better prompts

If a test fails, your AI interface needs work, not your code.

Agent Skills — First-Class Support

pytest-skill-engineering provides full Agent Skills spec compliance:

Compatibility field — Mark required tools, models, or platforms
Metadata — Title, description, version, attribution
Allowed-tools — Restrict which tools the agent can use
Scripts & Assets — Package Python scripts, prompts, and resources
Eval Bridge — Import evals from evals/evals.json, export grading results

Agent Skills are loaded natively when testing with CopilotEval — exactly as users experience them.

AI-Powered Reports

AI analyzes your results and tells you what to fix: which configuration to deploy, how to improve tool descriptions, where to cut costs. See a sample report →

Quick Start

# Install
uv add pytest-skill-engineering

# Authenticate (one-time)
gh auth login

# Run tests
pytest tests/

Configure AI Analysis (optional but recommended)

The AI-powered report needs a model to generate insights. Configure it in pyproject.toml:

[tool.pytest.ini_options]
addopts = "--aitest-summary-model=copilot/gpt-5-mini"

You can also use Azure OpenAI or other providers if you prefer — see Configuration.

Features

MCP Server Testing — Test tools, prompt templates, and bundled skills with real Copilot sessions
Agent Skills — Full agentskills.io spec compliance (compatibility, metadata, allowed-tools, evals bridge)
Custom Agents — Test .agent.md files and validate subagent dispatch
CLI Tool Testing — Verify Copilot can use command-line interfaces
Plugin Testing — Load complete plugin directories (plugin.json, .github/, .claude/ layouts) with auto-discovery
A/B Testing — Compare instructions, skills, custom agent versions, or tool configurations
Eval Leaderboard — Auto-ranked by pass rate and cost
Multi-Turn Sessions — Test conversations that build on context
Clarification Detection — Catch agents that ask questions instead of acting
LLM Assertions — Semantic checks with llm_assert, multi-dimension scoring with llm_score, image evaluation with llm_assert_image
AI-Powered Reports — Actionable feedback on tool descriptions, prompts, and costs
Cost Tracking — Copilot premium request tracking + USD estimation via pricing.toml

Who This Is For

MCP server authors — Validate that GitHub Copilot can actually use your tools
Agent Skills authors — Test skills exactly as users experience them in Copilot
Custom agent builders — Validate .agent.md instructions and subagent dispatch
Plugin developers — Test complete GitHub Copilot CLI plugins end-to-end
Teams shipping Copilot integrations — Catch skill stack regressions in CI/CD

Documentation

📚 Full Documentation

Requirements

Python 3.11+
pytest 9.0+
GitHub Copilot subscription (required)

Acknowledgments

Inspired by agent-benchmark.

License

MIT

Name		Name	Last commit message	Last commit date
Latest commit History 85 Commits
.copilot		.copilot
.github		.github
.squad		.squad
docs		docs
examples		examples
screenshots		screenshots
scripts		scripts
src/pytest_skill_engineering		src/pytest_skill_engineering
tests		tests
.env.example		.env.example
.gitattributes		.gitattributes
.gitignore		.gitignore
.pre-commit-config.yaml		.pre-commit-config.yaml
CHANGELOG.md		CHANGELOG.md
CODE_OF_CONDUCT.md		CODE_OF_CONDUCT.md
CONTRIBUTING.md		CONTRIBUTING.md
LICENSE		LICENSE
README.md		README.md
SECURITY.md		SECURITY.md
mkdocs.yml		mkdocs.yml
pricing.toml		pricing.toml
pyproject.toml		pyproject.toml
uv.lock		uv.lock

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Repository files navigation

pytest-skill-engineering

Why?

What This Tests

How It Works

Agent Skills — First-Class Support

AI-Powered Reports

Quick Start

Configure AI Analysis (optional but recommended)

Features

Who This Is For

Documentation

Requirements

Acknowledgments

License

About

Uh oh!

Releases 5

Packages

Uh oh!

Uh oh!

Contributors

Uh oh!

Languages

Folders and files

Latest commit

History

Repository files navigation

pytest-skill-engineering

Why?

What This Tests

How It Works

Agent Skills — First-Class Support

AI-Powered Reports

Quick Start

Configure AI Analysis (optional but recommended)

Features

Who This Is For

Documentation

Requirements

Acknowledgments

License

About

Resources

License

Code of conduct

Contributing

Security policy

Uh oh!

Stars

Watchers

Forks

Releases 5

Packages 0

Uh oh!

Uh oh!

Contributors

Uh oh!

Languages

Packages