Back to List
Notice:This resource is provided by a third-party author. Please review the code with AI tools or manually before use to ensure security and compatibility.
Pythonrobocurve/inspect-robots

inspect-robots

Open source evals for physical AI. Run any LLM/VLA on any arm/humanoid against any real/sim benchmark.

79.5/100
570Forks: 62
View on GitHubHomepage →
Loading report...

Similar Projects

opencompass

88

OpenCompass is an LLM evaluation platform, supporting a wide range of models from OpenAI, Anthropic, Gemini, Qwen, GLM, DeepSeek, etc, across 200+ datasets covering knowledge, reasoning, coding, science, language, long-context, and safety.

Python7.5K

ClawBench

81

Open-source benchmark for browser AI agents on daily tasks.

Python805

LightCompress

56

[EMNLP 2024 & AAAI 2026] A powerful toolkit for compressing large models including LLMs, VLMs, and video generative models.

Python749

ParseBench

78

ParseBench - A Document Parsing Benchmark for AI Agents

Python589
Back to List