AI & ML interests
None defined yet.
Recent Activity
View all activity
Papers
ClawsBench: Evaluating Capability and Safety of LLM Productivity Agents in Simulated Workspaces
SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks
datasets 25
benchflow/frontierphysics-pr889-evidence
Updated • 51
benchflow/frontierphysics-pr887-evidence
Updated • 63
benchflow/frontierphysics-pr888-evidence
Updated • 80
benchflow/frontierphysics-pr885-evidence
Updated • 68
benchflow/frontierphysics-pr884-evidence
Updated • 68
benchflow/frontierphysics-pr883-evidence
Updated • 67
benchflow/frontierphysics-pr881-evidence
Updated • 66
benchflow/frontierphysics-pr882-evidence
Updated • 104
benchflow/frontierphysics-pr879-evidence
Updated • 71
benchflow/frontierphysics-pr880-evidence
Updated • 59