OPEN SOURCE, OPEN TO EVERYONE

Open source. Open possibilities.

Discover quality open-source projects, submit projects anonymously, and claim and edit your own project.

Human-curated · Discover open source4092discovered

A little curiosity. A world of open source.

THE FIRST COLLECTION
Topic: llm-evaluation清除
sachncssachncs
ADDED

GitHub profile README for Sachin, a self-described applied AI architect and engineer. It points to selected open-source projects covering agents, retrieval (RAG), evaluation and reliability tooling, and lists skills, working style, engagement terms and contact details.

AI & MLDeveloper toolsAI agents
agentleakyagobski
ADDED

AgentLeak is an open-source privacy testing tool for AI agents. It detects data leaks across tool calls, memory, and logs using a local Python SDK, CLI, and MCP tools, providing redacted reports and CI gates without cloud dependencies.

Developer toolsAI & MLTesting & debugging
mlflowmlflow
ADDED

MLflow is an open-source AI engineering platform designed for managing the lifecycle of agents, LLMs, and traditional ML models, focusing on observability, evaluation, and deployment.

Film, Video & MediaVideo editing
sumOtotaO
ADDED

SUM is an open-source tool for verifying AI text transformations with cryptographic receipts. It checks what changed, what was preserved, and what was lost when AI rewrites text, with cross-runtime Ed25519 signatures and offline verification.

AI & MLDeveloper toolsAI assistants
garakNVIDIA
ADDED

garak is an open-source LLM vulnerability scanner from NVIDIA that probes generative AI models for hallucination, data leakage, prompt injection, toxicity, jailbreaks and other weaknesses, similar in spirit to nmap or Metasploit but for LLMs.

AI & MLDeveloper toolsTraining & evaluation
ifixaiifixai-ai
ADDED

iFixAi is an independent auditing tool for AI agents that evaluates whether an agent adheres to business KPIs and organizational structures through a structured scoring system.

AI & MLTraining & evaluation
deepevalconfident-ai
ADDED

DeepEval is an open-source LLM evaluation framework designed for unit testing AI applications, including RAG pipelines, agents, and chatbots.

AI & MLDeveloper toolsTraining & evaluation
ADDED

Open LLM hallucination benchmark regarding the BNCC (Brazil's National Common Curricular Base). It measures how much models invent official codes and texts using an auditable methodology with raw data and CI.

AI & MLTraining & evaluation