A benchmark that challenges language models to code solutions for scientific problems
-
Updated
Aug 3, 2026 - Python
A benchmark that challenges language models to code solutions for scientific problems
This document curates open-source projects, academic papers, capability benchmarks, and commercial solutions (international & China) in AI penetration testing, LLM red teaming, autonomous offensive agents, and vulnerability discovery—aimed at helping researchers, security engineers, and enterprise decision-makers quickly form a holistic view.
Agentic AI research papers, benchmarks, frameworks, and tools curated across 24 domains.
A quick view of high-performance convolution neural networks (CNNs) inference engines on mobile devices.
AI coding models, agents, CLIs, IDEs, AI app builders, open source tooling, benchmarks
Benchmark evaluating LLMs on their ability to create and resist disinformation. Includes comprehensive testing across major models (Claude, GPT-4, Gemini, Llama, etc.) with standardized evaluation metrics.
[NeurIPS 2025] AGI-Elo: How Far Are We From Mastering A Task?
A curated list of evaluation tools, benchmark datasets, leaderboards, frameworks, and resources for assessing model performance.
Live, open-source benchmark for comparing AI coding agents on real GitHub issues
The open benchmark for measuring the reliability, reproducibility, and determinism of AI context systems. Transparent metrics, reproducible tests, vendor-neutral results. Context Trust Levels CTL 0-4.
A measurement framework for autonomous AI agency across 7 dimensions · Paper: arXiv:2607.17947
Poolside Laguna S 2.1: Run 1M Context Locally (Tested) - Complete overview, benchmarks, local setup guides (vLLM, SGLang, llama.cpp), and test suite for Laguna S 2.1 118B MoE.
Live index of LLM evaluation tools and benchmarks, refreshed every 15 minutes from GitHub
An evidence-hostile, container-isolated benchmark and behavior analysis platform for long-horizon AI software agents.
Qualitative benchmark suite for evaluating AI coding agents and orchestration paradigms on realistic, complex development tasks
Pluralistic human evaluation infrastructure for AI in production
119 AI models × 55 benchmarks with per-score freshness dates, auto-updated pricing, task routing. Every score has a date and source URL. Daily CI.
Benchmarks AI conversations by estimated biochemical impact. Maps conversational outputs to neurochemical response profiles (oxytocin, dopamine, serotonin, endorphins, cortisol) via LLM analysis for physiological evaluation metrics.
A curated collection of AI model benchmarks and leaderboards — covering general rankings, coding, agents, reasoning, embeddings, and more
Independent benchmarks of AI capability on legal analysis tasks.
Add a description, image, and links to the ai-benchmarks topic page so that developers can more easily learn about it.
To associate your repository with the ai-benchmarks topic, visit your repo's landing page and select "manage topics."