Arena Testing
Status
- Available Now - PromptArena is actively being used for LLM testing
Overview
PromptArena (promptarena) is a CLI tool for running scenarios across multiple LLM providers, validating agent behavior, and generating comprehensive test reports.
PromptArena enables systematic testing of agent behavior — including multi-prompt routing, multi-turn conversations, and agent loops — with support for:
- Multi-Provider Testing - Run the same tests across OpenAI, Anthropic, Google, and more
- Scenarios & Self-Play - Multi-turn flows, agent-loop runs, and configurable role simulation
- Red-Team Mode - Adversarial scenarios for safety and robustness testing
- Multimodal Testing - Test with images, audio, and video content
- Comprehensive Reporting - HTML, JSON, JUnit XML, and Markdown reports
- Mock Testing - Fast, cost-free testing with mock providers
- CI/CD Integration - Built for automated testing pipelines
Installation
PromptArena is available as part of the PromptKit toolkit. See the PromptKit repository for installation instructions.
Quick Start
Basic Usage
# Run all tests with default configurationpromptarena run
# Specify configuration filepromptarena run --config my-arena.yaml
# Run specific providers onlypromptarena run --provider openai,anthropic
# Run specific scenariospromptarena run --scenario basic-qa,edge-casesConfiguration File
Create an arena.yaml configuration file:
apiVersion: promptkit.altairalabs.ai/v1alpha1kind: Arenametadata: name: my-arenaspec: prompt_configs: - id: assistant file: prompts/assistant.yaml
providers: - file: providers/openai.yaml - file: providers/anthropic.yaml
scenarios: - file: scenarios/test.yaml
defaults: output: dir: out formats: ["json", "html"]Key Features
Multi-Provider Testing
Test the same scenarios across multiple LLM providers:
# Compare OpenAI, Anthropic, and Geminipromptarena run --provider openai,anthropic,gemini --format htmlSupported providers include:
- OpenAI (GPT-4, GPT-3.5, etc.)
- Anthropic (Claude 3 Opus, Sonnet, Haiku)
- Google (Gemini Pro)
- Azure OpenAI
- And more
Multi-Turn Conversations
Define complex conversation flows in scenario files:
apiVersion: promptkit.altairalabs.ai/v1alpha1kind: Scenariometadata: name: customer-supportspec: task_type: support turns: - role: user parts: - type: text text: "I need help with my account" - role: assistant # Expected assistant response - role: user parts: - type: text text: "Can you reset my password?"Multimodal Content
Test with images, audio, and video:
turns: - role: user parts: - type: text text: "What's in this image?" - type: image media: file_path: test-data/sample.jpg detail: highSupported media types:
- Images: JPEG, PNG, GIF, WebP
- Audio: MP3, WAV, OGG, M4A
- Video: MP4, WebM, MOV
Self-Play Mode
Simulate realistic conversations with configurable roles:
# Enable self-play testingpromptarena run --selfplay
# Self-play with specific rolespromptarena run --selfplay --roles frustrated-customer,tech-supportMock Testing
Fast, cost-free testing during development:
# Use mock provider instead of real APIspromptarena run --mock-provider
# Use custom mock configurationpromptarena run --mock-config mock-responses.yamlOutput Formats
PromptArena generates comprehensive reports in multiple formats:
HTML Reports
Interactive HTML reports with:
- Side-by-side provider comparison
- Response times and token usage
- Cost analysis
- Media content visualization
- Test assertions pass/fail status
promptarena run --format htmlopen out/report-[timestamp].htmlJUnit XML
Standard JUnit XML for CI/CD integration:
promptarena run --format junit --junit-file out/junit.xmlProperties include:
media.images.total- Count of images testedmedia.loaded.success- Successfully loaded mediamedia.loaded.errors- Failed media loads- Test pass/fail status
JSON Reports
Machine-readable JSON for programmatic analysis:
promptarena run --format jsonMarkdown Reports
Human-readable markdown summaries:
promptarena run --format markdown --markdown-file out/results.mdTest Assertions
PromptArena supports multiple assertion types:
Text Assertions
assertions: - type: content_includes patterns: ["expected text", "another phrase"]
- type: content_excludes patterns: ["unwanted text"]Media Assertions
Validate media outputs from generative models:
# Image validationassertions: - type: image_format params: formats: [png, jpeg]
- type: image_dimensions params: width: 1920 height: 1080
# Audio validationassertions: - type: audio_format params: formats: [mp3, wav]
- type: audio_duration params: min_seconds: 29 max_seconds: 31
# Video validationassertions: - type: video_resolution params: presets: [4k, uhd]
- type: video_duration params: min_seconds: 59 max_seconds: 61CI/CD Integration
GitHub Actions
name: Arena Testson: [push, pull_request]
jobs: test: runs-on: ubuntu-latest steps: - uses: actions/checkout@v4
- name: Run Arena Tests run: promptarena run --ci --format junit env: OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }} ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}
- name: Publish Test Results uses: dorny/test-reporter@v1 if: always() with: name: Arena Test Results path: out/junit.xml reporter: java-junitCI Mode
Run in headless mode optimized for CI pipelines:
# Headless mode for CI pipelinespromptarena run --ci --format junit,jsonExit codes:
0- Success, all tests passed1- Failure, one or more tests failed
Commands
promptarena run
Run conversation simulations across multiple LLM providers.
Flags:
-c, --config- Configuration file path (default:arena.yaml)-j, --concurrency- Number of concurrent workers (default:6)--provider- Providers to use (comma-separated)--scenario- Scenarios to run (comma-separated)--temperature- Override temperature for all scenarios--max-tokens- Override max tokens for all scenarios--selfplay- Enable self-play mode--mock-provider- Replace all providers with MockProvider-o, --out- Output directory (default:out)--format- Output formats:json,junit,html,markdown-v, --verbose- Enable verbose debug logging
promptarena config-inspect
Inspect and validate arena configuration:
# Inspect default configurationpromptarena config-inspect
# Verbose output with detailspromptarena config-inspect --verbose
# JSON output for programmatic usepromptarena config-inspect --format jsonpromptarena prompt-debug
Test prompt generation with specific contexts:
# Test prompt generation for task typepromptarena prompt-debug --task-type support
# Test with regionpromptarena prompt-debug --task-type support --region us
# Test with scenario filepromptarena prompt-debug --scenario scenarios/customer-support.yamlpromptarena render
Generate HTML report from existing test results:
# Render from default locationpromptarena render out/index.json
# Custom output pathpromptarena render out/index.json --output custom-report.htmlBest Practices
Performance
# Increase concurrency for faster executionpromptarena run --concurrency 10
# Reduce concurrency for stabilitypromptarena run --concurrency 1Cost Control
# Use mock provider during developmentpromptarena run --mock-provider
# Test with cheaper models firstpromptarena run --provider gpt-3.5-turboReproducibility
# Always use same seed for consistent resultspromptarena run --seed 42Debugging
# Always start with config validationpromptarena config-inspect --verbose
# Use verbose mode to see API callspromptarena run --verbose --scenario problematic-test
# Test prompt generation separatelypromptarena prompt-debug --scenario scenarios/test.yamlLearn More
- Complete CLI Reference: Arena User Guide
- Configuration Reference: Config Documentation
- Getting Started: First Project Walkthrough
- CI/CD Integration: Pipeline Setup Guide
- GitHub Repository: AltairaLabs/PromptKit
Support
For questions, issues, or feature requests:
- Issues: GitHub Issues
- Discussions: GitHub Discussions