Public benchmarks leak
Once an evaluation is public, its shelf life is short. Scores become difficult to separate from memorization.
Neranu is the production system for creating, validating, running, and distributing interactive environments for AI training and evaluation.
$ pytest -q hidden_tests/
············FF
12 passed, 2 failed
$ agent: inspecting race condition_12 of 14 gates passed
Agents need worlds they can inspect, change, fail in, reset, and try again. Building those worlds is still a slow, manual craft.
Once an evaluation is public, its shelf life is short. Scores become difficult to separate from memorization.
Weak verifiers produce impressive metrics without proving that the agent solved the intended task.
Mutable dependencies, hidden state, and missing provenance make results impossible to trust or replay.
Every environment moves through an explicit lifecycle. Each gate produces evidence, has an owner, and is attached to an immutable revision.
Explore the validation engineCreate environments, prove their quality, freeze releases, run models, and license the result from one connected system.
Turn repositories, problems, and workflows into versioned environments with runtime, reward, and provenance in one workspace.
Test reproducibility, reward integrity, hidden checks, security boundaries, and resistance to shortcuts before release.
A system of record for every environment, revision, artifact, right, release, and downstream use.
Compare model snapshots, inspect trajectories, cluster failures, and trace every result to a frozen environment revision.
Commission verified packs and license environments from qualified domain experts and data owners.
Environment release 1.4.0
A Neranu environment is more than a prompt and a container. The specification, verifier, rights, reviews, baselines, and exact runtime are frozen into a portable release record.
Sealed, refreshable eval packs built around the failure surface of your next model release.
Build a private evalExecutable tasks with deterministic rewards, repeatable reset, and complete lineage.
Commission a collectionVersioned repositories, realistic issues, hidden tests, and reproducible agent workspaces.
In developmentPersistent tool-use and multi-step environments served through a unified rollout API.
On the roadmapThe same production system that verifies an answer or patch can govern a repository, a tool workflow, or a long-running agent world.
Explore the product through a realistic sample workspace.