Repository: vibrantlabsai/ragas
Stars: 13436
CLAUDE.md
CLAUDE.md
This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.
Project Overview
Ragas is an evaluation toolkit for Large Language Model (LLM) applications. It provides objective metrics for evaluating LLM applications, test data generation capabilities, and integrations with popular LLM frameworks.
The repository contains:
1. Ragas Library - The main evaluation toolkit including experimental features (in src/ragas/ directory)
- Core evaluation metrics and test generation
- Experimental features available at ragas.experimental
Development Environment Setup
Installation
Choose the appropriate installation based on your needs:
RECOMMENDED: Minimal dev setup (79 packages - fast)
make install-minimalFULL: Complete dev environment (383 packages - comprehensive)
make installOR manual installation:
Create a virtual environment
python -m venv venv
source venv/bin/activate # On Windows, use venv\Scripts\activateMinimal dev setup (uses [project.optional-dependencies].dev-minimal)
uv pip install -e ".[dev-minimal]"Full dev setup (uses [dependency-groups].dev)
uv sync --group devInstallation Methods Explained
- Minimal setup: Uses uv pip install with optional dependencies for selective installation
- Full setup: Uses uv sync with dependency groups for comprehensive environment management
- No naming conflicts: dev-minimal vs dev clearly distinguish the two approaches
Workspace Structure
The project uses a UV workspace configuration for managing multiple packages:
Install
uv syncInstall examples separately
uv sync --package ragas-examplesBuild specific workspace package
uv build --package ragas-examplesWorkspace Members:
- ragas (main package) - Located in src/ragas/
- ragas-examples (examples package) - Located in examples/
The workspace ensures consistent dependency versions across packages and enables editable installs of workspace members.
Common Commands
Commands (from root directory)
Setup and installation
make install-minimal # Minimal dev setup (79 packages - recommended)
make install # Full dev environment (383 packages - complete)Code quality
make format # Format and lint all code
make type # Type check all code
make check # Quick health check (format + type, no tests)Testing
make test # Run all unit tests
make test-e2e # Run end-to-end testsCI/Build
make run-ci # Run complete CI pipeline
make clean # Clean all generated filesDocumentation
make build-docs # Build all documentation
make serve-docs # Serve documentation locallyBenchmarks
make benchmarks # Run performance benchmarks
make benchmarks-docker # Run benchmarks in DockerTesting
Run all tests (from root)
make testRun specific test (using pytest -k flag)
make test k="test_name"Run end-to-end tests
make test-e2eDirect pytest commands for more control
uv run pytest tests/unit -k "test_name"
uv run pytest tests/unit -vDocumentation
Build all documentation (from root)
make build-docsServe documentation locally
make serve-docsBenchmarks
Run all benchmarks locally
make benchmarksRun benchmarks in Docker
make benchmarks-dockerProject Architecture
The repository has the following structure:
/ # Main ragas project
├── src/ragas/ # Source code including experimental features
│ └── experimental/ # Experimental features
├── tests/ # All tests (core + experimental)
│ └── experimental/ # Experimental tests
├── examples/ # Example code
├── pyproject.toml # Build config
├── docs/ # Documentation
├── scripts/ # Build/CI scripts
├── Makefile # Build commands
└── README.md # Repository overviewRagas Core Components
The Ragas core library provides metrics, test data generation and evaluation functionality for LLM applications:
1. Metrics - Various metrics for evaluating LLM applications including:
- AspectCritic
- AnswerCorrectness
- ContextPrecision
- ContextRecall
- Faithfulness
- and many more
2. Test Data Generation - Automatic creation of test datasets for LLM applications
3. Integrations - Integrations with popular LLM frameworks like LangChain, LlamaIndex, and observability tools
Experimental Components
The experimental features are now integrated into the main ragas package:
1. Experimental features are available at ragas.experimental
2. Dataset and Experiment management - Enhanced data handling for experiments
3. Advanced metrics - Extended metric capabilities
4. Backend support - Multiple storage backends (CSV, JSONL, Google Drive, in-memory)
To use experimental features:
from ragas import Dataset
from ragas import experiment
from ragas.backends import get_registryDebugging Logs
To view debug logs for any module:
import loggingConfigure logging for a specific module (example with analytics)
analytics_logger = logging.getLogger('ragas._analytics')
analytics_logger.setLevel(logging.DEBUG)Create a console handler and set its level
console_handler = logging.StreamHandler()
console_handler.setLevel(logging.DEBUG)Create a formatter and add it to the handler
formatter = logging.Formatter('%(name)s - %(levelname)s - %(message)s')
console_handler.setFormatter(formatter)Add the handler to the logger
analytics_logger.addHandler(console_handler)Memories
- whenever you create such docs put in in /\_experiments because that is gitignored and you can use it as a scratchpad or tmp directory for storing these
- always use uv to run python and python related commandline tools like isort, ruff, pyright etc. This is because we are using uv to manage the .venv and dependencies.
- The project uses two distinct dependency management approaches:
- Minimal setup: [project.optional-dependencies].dev-minimal for fast development (79 packages)
- Full setup: [dependency-groups].dev for comprehensive development (383 packages)
- Use make install-minimal for most development tasks, make install for full ML stack work
- if the user asks you to save a plan, save it into the plan/ directory with an appropriate file name.
README.md
<h1 align="center">
<img style="vertical-align:middle" height="200"
src="https://raw.githubusercontent.com/vibrantlabsai/ragas/main/docs/_static/imgs/logo.png">
</h1>
<p align="center">
<i>Supercharge Your LLM Application Evaluations 🚀</i>
</p>
<p align="center">
<a href="https://github.com/vibrantlabsai/ragas/releases">
<img alt="Latest release" src="https://img.shields.io/github/release/vibrantlabsai/ragas.svg">
</a>
<a href="https://www.python.org/">
<img alt="Made with Python" src="https://img.shields.io/badge/Made%20with-Python-1f425f.svg?color=purple">
</a>
<a href="https://github.com/vibrantlabsai/ragas/blob/master/LICENSE">
<img alt="License Apache-2.0" src="https://img.shields.io/github/license/vibrantlabsai/ragas.svg?color=green">
</a>
<a href="https://pypi.org/project/ragas/">
<img alt="Ragas Downloads per month" src="https://static.pepy.tech/badge/ragas/month">
</a>
<a href="https://discord.gg/5djav8GGNZ">
<img alt="Join Ragas community on Discord" src="https://img.shields.io/discord/1119637219561451644">
</a>
<a target="_blank" href="https://deepwiki.com/vibrantlabsai/ragas">
<img
src="https://devin.ai/assets/deepwiki-badge.png"
alt="Ask DeepWiki.com"
height="20"
/>
</a>
</p>
<h4 align="center">
<p>
<a href="https://docs.ragas.io/">Documentation</a> |
<a href="#fire-quickstart">Quick start</a> |
<a href="https://discord.gg/5djav8GGNZ">Join Discord</a> |
<a href="https://blog.ragas.io/">Blog</a> |
<a href="https://newsletter.ragas.io/">NewsLetter</a> |
<a href="https://www.ragas.io/careers">Careers</a>
<p>
</h4>
Objective metrics, intelligent test generation, and data-driven insights for LLM apps
Ragas is your ultimate toolkit for evaluating and optimizing Large Language Model (LLM) applications. Say goodbye to time-consuming, subjective assessments and hello to data-driven, efficient evaluation workflows.
Don't have a test dataset ready? We also do production-aligned test set generation.
Key Features
- 🎯 Objective Metrics: Evaluate your LLM applications with precision using both LLM-based and traditional metrics.
- 🧪 Test Data Generation: Automatically create comprehensive test datasets covering a wide range of scenarios.
- 🔗 Seamless Integrations: Works flawlessly with popular LLM frameworks like LangChain and major observability tools.
- 📊 Build feedback loops: Leverage production data to continually improve your LLM applications.
:shield: Installation
Pypi:
pip install ragasAlternatively, from source:
pip install git+https://github.com/vibrantlabsai/ragas:fire: Quickstart
Clone a Complete Example Project
The fastest way to get started is to use the ragas quickstart command:
List available templates
ragas quickstartCreate a RAG evaluation project
ragas quickstart rag_evalSpecify where you want to create it.
ragas quickstart rag_eval -o ./my-projectAvailable templates:
- rag_eval - Evaluate RAG systems
Coming Soon:
- agent_evals - Evaluate AI agents
- benchmark_llm - Benchmark and compare LLMs
- prompt_evals - Evaluate prompt variations
- workflow_eval - Evaluate complex workflows
Evaluate your LLM App
ragas comes with pre-built metrics for common evaluation tasks. For example, Aspect Critique evaluates any aspect of your output using DiscreteMetric:
import asyncio
from openai import AsyncOpenAI
from ragas.metrics import DiscreteMetric
from ragas.llms import llm_factorySetup your LLM
client = AsyncOpenAI()
llm = llm_factory("gpt-4o", client=client)Create a custom aspect evaluator
metric = DiscreteMetric(
name="summary_accuracy",
allowed_values=["accurate", "inaccurate"],
prompt="""Evaluate if the summary is accurate and captures key information.Response: {response}
Answer with only 'accurate' or 'inaccurate'."""
)
Score your application's output
async def main():
score = await metric.ascore(
llm=llm,
response="The summary of the text is..."
)
print(f"Score: {score.value}") # 'accurate' or 'inaccurate'
print(f"Reason: {score.reason}")
if __name__ == "__main__":
asyncio.run(main())
Note: Make sure your OPENAI_API_KEY environment variable is set.Find the complete Quickstart Guide
Want help in improving your AI application using evals?
In the past 2 years, we have seen and helped improve many AI applications using evals. If you want help with improving and scaling up your AI application using evals.
🔗 Book a slot or drop us a line: [email protected].
🫂 Community
If you want to get more involved with Ragas, check out our discord server. It's a fun community where we geek out about LLM, Retrieval, Production issues, and more.
Contributors
+----------------------------------------------------------------------------+
| +----------------------------------------------------------------+ |
| | Developers: Those who built with ragas. | |
| | (You have import ragas somewhere in your project) | |
| | +----------------------------------------------------+ | |
| | | Contributors: Those who make ragas better. | | |
| | | (You make PR to this repo) | | |
| | +----------------------------------------------------+ | |
| +----------------------------------------------------------------+ |
+----------------------------------------------------------------------------+We welcome contributions from the community! Whether it's bug fixes, feature additions, or documentation improvements, your input is valuable.
1. Fork the repository
2. Create your feature branch (git checkout -b feature/AmazingFeature)
3. Commit your changes (git commit -m 'Add some AmazingFeature')
4. Push to the branch (git push origin feature/AmazingFeature)
5. Open a Pull Request
🔍 Open Analytics
At Ragas, we believe in transparency. We collect minimal, anonymized usage data to improve our product and guide our development efforts.
✅ No personal or company-identifying information
✅ Open-source data collection code
✅ Publicly available aggregated data
To opt-out, set the RAGAS_DO_NOT_TRACK environment variable to true.
Cite Us
@misc{ragas2024,
author = {VibrantLabs},
title = {Ragas: Supercharge Your LLM Application Evaluations},
year = {2024},
howpublished = {\url{https://github.com/vibrantlabsai/ragas}},
}