1.0
Data and evaluation infrastructure
for frontier AI.
AI systems are only as capable as the data they train on and the benchmarks they improve against. The teams building the most useful AI will win by having better training data and more precise measurement infrastructure than anyone else. That infrastructure is what General Machines builds.
We make datasets, evaluations, and benchmarks for frontier AI labs and applied teams. Our work covers two domains: the agentic web, where we measure and train AI agents operating across the internet, and physical AI, where we build the ego-centric and embodied datasets that teach robots and spatial systems how to perceive and act in the real world.
Collecting high-quality behavioral data at scale is expensive, unglamorous, and undervalued. We think it is also the most important bottleneck in AI progress. Frontier labs can train better models if they have better data. Applied teams can ship better agents if they can measure agent behavior precisely. We exist to close both gaps.
Our first applied research arm is Machine Commerce. It focuses entirely on agentic commerce: tracing how AI agents shop, compare, and transact online, then turning those behavioral traces into evaluations and benchmarks that help labs build trustworthy commercial agents. Every trace is collected in a live retail environment, not a sandbox.
General Machines is building the measurement and data infrastructure that the next era of AI depends on.