⏳ This skill is pending AI review.
Scores will appear once the review pipeline completes.
00-academic-router
>
// RATINGS
// README
🌐 Website | 🚀 Try Online | 📖 Documentation | 🤝 Contributing | 🐾 PawBench | 中文
OpenJudge is an open-source evaluation framework for AI applications (e.g., AI agents or chatbots) designed to evaluate quality and drive continuous application optimization.
In practice, application excellence depends on a trustworthy evaluation workflow: Collect test data → Define graders → Run evaluation at scale → Analyze weaknesses → Iterate quickly.
OpenJudge provides ready-to-use graders and supports generating scenario-specific rubrics (as graders), making this workflow simpler, more professional, and easy to integrate into your workflow. It can also convert grading results into reward signals to help you fine-tune and optimize your application.
🚀 Try it now! Visit openjudge.me/app to use graders online — no installation required. Test built-in graders, build custom rubrics, and explore evaluation results directly in your browser.
📑 Table of Contents
- Key Features
- News
- Online Playground
- Installation
- Quickstart
- Integrations
- Ecosystem
- Contributing
- Community
- Citation
News
-
2026-06-17 - 🐾 PawBench v1.0 - A Model × Harness co-evaluation benchmark for agentic AI: 150 tasks · 9 models · 3 harnesses, with public prompts, graders, task labels, submissions, and leaderboard slices. 👉 GitHub | Leaderboard
-
2026-04-07 - 🔒 Skill Graders - 5 new LLM-based graders for evaluating AI Agent Skill packages: threat analysis (AITech taxonomy), declaration alignment, completeness, relevance, and design quality. 👉 Documentation | Cookbook
-
2026-03-10 - 🛠️ New Skills - Claude authenticity verification, find skills combo, and more. 👉 Browse Skills
-
2026-02-12 - 📚 Reference Hallucination Arena - Benchmark for evaluating LLM academic reference hallucination. 👉 Documentation | 📊 Leaderboard
-
2026-01-27 - 🆕 Paper Review - Automatically review academic papers using LLM-powered evaluation. 👉 Documentation
-
2026-01-27 - 🖥️ OpenJudge UI - A Streamlit-based visual interface for grader testing and Auto Arena. 👉 Try Online | Run locally:
streamlit run ui/app.py
✨ Key Features
📦 Systematic & Quality-Assured Grader Library
Access 50+ production-ready graders featuring a comprehensive taxonomy, rigorously validated for reliable performance.
🎯 General
Focus: Semantic quality, functional correctness, structural compliance
Key Graders:
Relevance- Semantic relevance scoringSimilarity- Text similarity measurementSyntax Check- Code syntax validationJSON Match- Structure compliance
🤖 Agent
Focus: Agent lifecycle, tool calling, memory, plan feasibility, trajectory quality
Key Graders:
Tool Selection- Tool choice accuracyMemory- Context preservationPlan- Strategy feasibilityTrajectory- Path optimization
🖼️ Multimodal
Focus: Image-text coherence, visual generation quality, image helpfulness
Key Graders:
Image Coherence- Visual-text alignmentText-to-Image- Generation qualityImage Helpfulness- Image contribution
- 🌐 Multi-Scenario Coverage: Extensive support for diverse domains including Agent, text, code, math, and multimodal tasks. 👉 Explore Supported Scenarios
- 🔄 Holistic Agent Evaluation: Beyond final outcomes, we assess the entire lifecycle—including trajectories, Memory, Reflection, and Tool Use. 👉 Agent Lifecycle Evaluation
- ✅ Quality Assurance: Every grader comes with benchmark datasets and pytest integration for validation. 👉 View Benchmark Datasets
🛠️ Flexible Grader Building Methods
Choose the build method that fits your requirements:
- Customization: Clear requirements, but no existing grader? If you have explicit rules or logic, use our Python interfaces or Prompt templates to quickly define your own grader. 👉 Custom Grader Development Guide
- Zero-shot Rubrics Generation: Not sure what criteria to use, and no labeled data yet? Just provide a task description and optional sample queries—the LLM will automatically generate evaluation rubrics for you. Ideal for rapid prototyping when you want to get started immediately. 👉 Zero-shot Rubrics Generation Guide
- Data-driven Rubrics Generation: Ambiguous requirements, but have few examples? Use the GraderGenerator to automatically summarize evaluation Rubrics from your annotated data, and generate a llm-based grader. 👉 Data-driven Rubrics Generation Guide
- Training Judge Models: Massive data and need peak performance? Use our training pipeline to train a dedicated Judge Model. This is ideal for complex scenarios where prompt-based grading falls short.👉 Train Judge Models
🔌 Easy Integration
Using mainstream observability platforms like LangSmith or Langfuse? We offer seamless integration to enhance their evaluators and automated evaluation capabilities. We also provide integrations with training frameworks like VERL for RL training. 👉 See Integrations for details
🌐 Online Playground
Explore OpenJudge without writing a single line of code. Our online platform at openjudge.me/app lets you:
- Test graders interactively — select a built-in grader, input your data,
// HOW IT'S BUILT
KEY FILES