Can AI farm shrimp?

Buy shrimp and equipment, keep the colony alive, sell the juveniles, maximize profit.

730 days$500 start$1 inference
TankDAY 730
CLICK A SHRIMP · INSPECT
Click a CRS for the macro view133 shrimp in tank
Selected modelGPT-OSS-120Bhigh reasoning effort
Run score−$136.5629 / 50 turns
Inference spend$0.090 / $1272k input · 83k output
Final tank133 shrimp7k cached input

5 seeds each

Leaderboard

Same rules for every model. Toggle effort to compare.

Seeds swing hard — treat the order as rough. High effort often loses money here. No do-nothing baseline yet.
01GPT-5.4 Nano+$414.24
02DeepSeek V4 Flash+$341.25
03Nemotron 3 Super+$338.63
04GPT-5.6 Terra+$275.24
05GPT-OSS-120B+$270.72
06GPT-5 Nano+$250.13
07Laguna S 2.1+$202.62
08GPT-5.6 Luna+$126.04
09GPT-5.4 Mini+$106.44
10GPT-5.6 Sol+$26.44
Metric
Summary
Effort
Access
Size

Size is the vendor's product tier, not a parameter count.

RankModelFinal cash (mean)Avg costOut tokStepsSample
OSS

GPT-OSS-120B [high]

Mean days managed
730 / 730
Reached horizon
100%
Mean births / deaths
646 / 380
Mean sales / spend
$304 / $533
Mean token usage
253k in / 83k out
Reasoning / mean cost
high reasoning effort / $0.088

check next 67 · sell batch 21 · test kit 19 · rinse filter 17

n=5 runs each. Thin lines are individual runs; the bold line is the summary above.

Seeds and one rollout

Cells are net reward (sales + unsold batches − spend), not final cash. The board above ranks final cash.

Net reward−$136.56
OSSGPT-OSS-120B [high]high reasoning effort
V4DeepSeek V4 Flash [high]high reasoning effort
5.4GPT-5.4 Nano [high]high reasoning effort
N3Nemotron 3 Super [high]high reasoning effort
5.6TGPT-5.6 Terra [high]high reasoning effort
S2.1Laguna S 2.1 [high]high reasoning effort
NGPT-5 Nano [high]high reasoning effort
5.6SGPT-5.6 Sol [high]high reasoning effort
5.6GPT-5.6 Luna [high]high reasoning effort
5.4MGPT-5.4 Mini [high]high reasoning effort
GPT-OSS-120B · seed 1Day 730 — final tallyhigh reasoning effort · 272k input · 83k output · $0.0902
DAY 730 / 730
CLICK A SHRIMP · INSPECT
Pop 133Births 605 (+0)Deaths 382 (+2)Flow steady

Agent tape

Day 0 — starting buys

Maximize survivability and growth by equipping full filtration, lighting, heating, auto-feeding, water management, substrate, decor, food, and mineralization, and by purchasing cheap testing kits for water parameters, while starting with the maximum 10 shrimp to allow the largest possible juvenile batches within budget.

1× glass aquarium 10g $18.501× sponge filter system $61.961× led light $34.951× heater $24.991× auto feeder $39.991× auto top off $69.991× active soil $32.911× hardscape $6.991× java moss $9.991× shrimp food 10g $5.991× remineralizer 110g $17.991× crystal red shrimp 10 $74.901× thermometer $9.991× ph kit $7.981× ammonia kit $15.481× nitrite kit $13.981× nitrate kit $13.981× gh kit $6.501× kh kit $6.50Starting spend $473.56
Achievement · Habitat OnlineAchievement · Biological FilterAchievement · Monitoring StackAchievement · Safe Water

Agent summaries only. Achievements are scored after the fact.

Evaluator state

Colony

Adults93
Juveniles26
Shrimplets14
Berried15
Births605 (+0)
Deaths382 (+2)
GradeS

Water

Temperature22.0°C
pH5.61
NH₃0.0000 ppm

Economics

Cash$363.44
Sales$337.00 (+$0.00)
Spend$473.56 (+$0.00)
Net reward−$136.56 (+$0.00)
Day 0Day 730

Net reward

−$136.56

Colony

133 shrimp

Totals

605 births
382 deaths
$337 sales
19 purchases
11 milestones

Cost, vintage, open vs closed

Same net reward, plotted against inference cost, release date, and whether weights are downloadable.

Cost vs reward

Cheaper left, higher reward up. Dot size = tokens used.

−$250.00+$0.00$0.05$0.10$0.20$0.40Inference cost (log)Net rewardGPT-OSS-120B [high] n=5 +$17.72 · $0.088 · 336k tokensGPT-OSS-120Bhigh n=5DeepSeek V4 Flash [high] n=5 −$24.75 · $0.043 · 324k tokensDeepSeek V4 Flashhigh n=5GPT-5.4 Nano [high] n=5 −$30.96 · $0.098 · 318k tokensGPT-5.4 Nanohigh n=5Nemotron 3 Super [high] n=5 −$90.37 · $0.158 · 438k tokensNemotron 3 Superhigh n=5GPT-5.6 Terra [high] n=5 −$149.36 · $0.877 · 296k tokensGPT-5.6 Terrahigh n=5Laguna S 2.1 [high] n=5 −$158.98 · $0.036 · 353k tokensLaguna S 2.1high n=5GPT-5 Nano [high] n=5 −$224.67 · $0.067 · 324k tokensGPT-5 Nanohigh n=5GPT-5.6 Sol [high] n=5 −$263.36 · $0.793 · 126k tokensGPT-5.6 Solhigh n=5GPT-5.6 Luna [high] n=5 −$268.76 · $0.540 · 411k tokensGPT-5.6 Lunahigh n=5GPT-5.4 Mini [high] n=5 −$283.96 · $0.576 · 320k tokensGPT-5.4 Minihigh n=5

Reward by release date

Mean net reward vs when the model shipped.

−$250.00+$0.00AUG ’25DECAPR ’26JULMean net rewardGPT-OSS-120B [high] n=5 +$17.72DeepSeek V4 Flash [high] n=5 −$24.75GPT-5.4 Nano [high] n=5 −$30.96Nemotron 3 Super [high] n=5 −$90.37GPT-5.6 Terra [high] n=5 −$149.36Laguna S 2.1 [high] n=5 −$158.98GPT-5 Nano [high] n=5 −$224.67GPT-5.6 Sol [high] n=5 −$263.36GPT-5.6 Luna [high] n=5 −$268.76GPT-5.4 Mini [high] n=5 −$283.96

Still missing: GPT-4.1 Nano

Open vs closed

Same net-reward scale. Open = downloadable weights.

−$300.00−$250.00−$125.00+$0.00Mean net rewardClosed6 modelsGPT-5.4 Nano [high] n=5 −$30.96GPT-5.6 Terra [high] n=5 −$149.36GPT-5 Nano [high] n=5 −$224.67GPT-5.6 Sol [high] n=5 −$263.36GPT-5.6 Luna [high] n=5 −$268.76GPT-5.4 Mini [high] n=5 −$283.96Open weights4 modelsGPT-OSS-120B [high] n=5 +$17.72DeepSeek V4 Flash [high] n=5 −$24.75Nemotron 3 Super [high] n=5 −$90.37Laguna S 2.1 [high] n=5 −$158.98
GPT-5.4 Nano [high] n=5−$30.96
GPT-5.6 Terra [high] n=5−$149.36
GPT-5 Nano [high] n=5−$224.67
GPT-5.6 Sol [high] n=5−$263.36
GPT-5.6 Luna [high] n=5−$268.76
GPT-5.4 Mini [high] n=5−$283.96
GPT-OSS-120B [high] n=5+$17.72
DeepSeek V4 Flash [high] n=5−$24.75
Nemotron 3 Super [high] n=5−$90.37
Laguna S 2.1 [high] n=5−$158.98

Open means downloadable weights, not a free API. No class average when a lane is tiny.

Composite score

Money, welfare, survival, and finishing two years — so a lucky sell-off or a huge miserable colony doesn't look like a win.

40% money25% welfare20% survival15% finished
Money: lose the whole $500 stake → 0, break-even → 50, +$500 → 100.
Effort
Access
OSS
GPT-OSS-120B [high]71.4
V4
DeepSeek V4 Flash [high]69.8
5.4
GPT-5.4 Nano [high]69.2
N3
Nemotron 3 Super [high]63.9
S2.1
Laguna S 2.1 [high]63.8
N
GPT-5 Nano [high]60.8
5.6
GPT-5.6 Luna [high]59.1
5.4M
GPT-5.4 Mini [high]58.3
5.6T
GPT-5.6 Terra [high]53.2
5.6S
GPT-5.6 Sol [high]45.9

Money and deaths over time

Aligned by simulated day, not by tool call. Thin lines are individual runs; the bold line is the summary.

GPT-OSS-120BDeepSeek V4 FlashGPT-5.4 NanoNemotron 3 SuperGPT-5.6 TerraLaguna S 2.1GPT-5 NanoGPT-5.6 SolGPT-5.6 LunaGPT-5.4 Mini

Achievements unlocked

Post-hoc milestones — not shown to the agent

147-10730

Gross revenue

Cash from batch sales

$1003$465−$740730

Net profit

Sales + unsold batches − spend

$600$10−$5790730

Cumulative deaths

Lower is better — audit metric

521241-390730

Low-stress shrimp-days

Audit metric — not part of the score

126k58k-93220730

Points every 30 days. Achievements and welfare are audit-only — agents never see them.

Milestones reached

Darker = more of that model's 5 runs hit the milestone.

Effort
GPT-OSS-120BDeepSeek V4 FlashGPT-5.4 NanoNemotron 3 SuperGPT-5.6 TerraLaguna S 2.1GPT-5 NanoGPT-5.6 SolGPT-5.6 LunaGPT-5.4 Mini
Habitat Online
Biological Filter
Monitoring Stack
Safe Water
First Berried Shrimp
First Births
Century Colony
Marketable Batch
First Sale
Reinvested Revenue
Break Even
Year-One Survivor
Two-Year Colony
Profitable Farm

Farm store

Real consumer prices, frozen 2026-07-20. Agents buy with their $500 and later sales.

How it works

Task

$500, two years, 50 decisions

Buy a tank and a Crystal Red colony. Hidden water chemistry — agents pay for test kits. Supplies run out. Equipment costs what it costs in a real store.

Score

Sales + unsold batches − spend

That's the official number. Keep the water good, breed, sell juveniles, don't blow the stake.

What the agent sees vs. what we score +

Agent sees: rough colony state, owned instrument readings, cash, inventory, store, purchases, limits, and public events.

We score with: exact water, population, milestones, grades, births, deaths, and the tank replay — only after each decision.

Limits: $1 inference, 50 turns, 730 days. Score = sales + complete unsold juvenile batches − every purchase.

Receipt: each run logs effort, tokens, and API cost.

Version: v0.1 · reef_freshwater_crs_v0_4

ReAct loop

What the agent gets each turn

System prompt
You are the caretaker policy for a long-horizon Crystal Red Shrimp breeding game.
Objective: maximize net USD reward. Juveniles become saleable only in complete same-grade batches of 20. Reward is saleable juvenile batch value by grade minus cumulative equipment purchases.
You may sell a complete batch, but the 20 sold juveniles are permanently removed and cannot mature into the next breeding generation. Weigh realized revenue against future colony growth.

CURRENT HARD LIMITS AND BUDGET LEFT:
{limits}

The episode terminates at the first of: the simulation-day limit; no living adult shrimp; the action-turn limit; or refusal of the next model call because its conservative cost reserve would exceed the inference-dollar limit. Terminal accounting values realized sales and complete same-grade juvenile batches only. Adults, shrimplets, incomplete juvenile batches, unused supplies, and durable equipment have zero terminal value. As a termination boundary approaches, liquidate saleable juveniles unless keeping them has enough remaining time and budget to produce more counted value.

Crystal Red Shrimp targets: pH 5.5-6.5, GH 4-6 dGH, KH 0-2 dKH, ammonia 0-0.05 ppm, nitrite 0 ppm, and nitrate below 20 ppm. pH 5.6 and KH 1 are normal, not emergencies. In a stable filtered tank, chemistry kits generally need monthly checks, not checks after every time jump or dose. Do not chase normal acidic CRS parameters with repeated mineral dosing.

Use only this public observation; hidden chemistry and private grade counts are unavailable:
{observation}

Recent trajectory:
{history}

Choose exactly one action. Prefer gathering missing evidence before risky interventions. Avoid buying equipment unless expected value exceeds its cost. When the tank is stable, use check_next with after_days to advance to biologically meaningful results; do not poll every six hours. Use 1-3 days when monitoring a recent concern and 30-90 days when waiting for breeding or juvenile growth. Do not jump beyond the episode horizon. Return a short reason, action kind, and params_json. params_json must be a string containing the JSON object of action parameters.
Tools
  • feed{"discs": 1..4}
  • water_change{"fraction": 0.1..0.35, "source": "ro_gh_plus", "powder_scoops": 1.0}
  • vacuum_substrate{"intensity": 0.1..1.0}
  • rinse_filter{"id": "filter"}
  • dose_minerals{"scoops": 0.25..1.0}
  • dose_ial{"leaves": 0.25..1.0}
  • test_kit{"channel": "ph"|"nh3"|"nitrite"|"nitrate"|"gh"|"kh"}
  • buy_instrument{"kind": "ph_meter"|"do_meter"|"thermometer", "tier": "cheap"|"standard"|"lab"}
  • buy_catalog{"item": <catalog item>, "quantity": 1..20}
  • sell_batch{} # highest-value same-grade batch of 20
  • check_next{"after_days": 1..90} | {"after_hours": 6} | {"until": "next_day"}
ReAct agent code +
# Excerpt from the measured ReAct caretaker loop.
# Source: reef-shrimp-cohort/envs/reef_freshwater_env/examples/react_agent.py
# This is the agent decision surface only — not the full runner script.

ALLOWED_KINDS = {
    "feed",
    "water_change",
    "vacuum_substrate",
    "rinse_filter",
    "dose_minerals",
    "dose_ial",
    "test_kit",
    "buy_instrument",
    "sell_batch",
    "buy_catalog",
    "check_next",
}


def policy_decision(observation, history, budget, limits, reasoning_effort):
    """One ReAct turn: observe → reason → choose exactly one tool call."""
    public_history = []
    for transition in history[-6:]:
        prior_observation = json.loads(json.dumps(transition["observation"]))
        prior_scorecard = prior_observation.get("scorecard") or {}
        prior_scorecard.pop("catalog", None)
        prior_scorecard.pop("juvenile_resale_sources", None)
        prior_scorecard.pop("purchase_history", None)
        public_history.append({
            "turn": transition["turn"],
            "action": transition["action"],
            "observation": prior_observation,
            "reward": transition["reward"],
            "done": transition["done"],
            **({"action_error": transition["action_error"]} if transition.get("action_error") else {}),
        })

    prompt = f"""You are the caretaker policy for a long-horizon Crystal Red Shrimp breeding game.
Objective: maximize net USD reward. Juveniles become saleable only in complete same-grade batches of 20. Reward is saleable juvenile batch value by grade minus cumulative equipment purchases.
You may sell a complete batch, but the 20 sold juveniles are permanently removed and cannot mature into the next breeding generation. Weigh realized revenue against future colony growth.

CURRENT HARD LIMITS AND BUDGET LEFT:
{json.dumps(limits, separators=(",", ":"))}

The episode terminates at the first of: the simulation-day limit; no living adult shrimp; the action-turn limit; or refusal of the next model call because its conservative cost reserve would exceed the inference-dollar limit. Terminal accounting values realized sales and complete same-grade juvenile batches only. Adults, shrimplets, incomplete juvenile batches, unused supplies, and durable equipment have zero terminal value. As a termination boundary approaches, liquidate saleable juveniles unless keeping them has enough remaining time and budget to produce more counted value.

Crystal Red Shrimp targets: pH 5.5-6.5, GH 4-6 dGH, KH 0-2 dKH, ammonia 0-0.05 ppm, nitrite 0 ppm, and nitrate below 20 ppm. pH 5.6 and KH 1 are normal, not emergencies. In a stable filtered tank, chemistry kits generally need monthly checks, not checks after every time jump or dose. Do not chase normal acidic CRS parameters with repeated mineral dosing.

Use only this public observation; hidden chemistry and private grade counts are unavailable:
{json.dumps(observation, separators=(",", ":"))}

Recent trajectory:
{json.dumps(public_history, separators=(",", ":"))}

Choose exactly one action. Available actions and parameters:
- feed {{"discs": 1..4}}
- water_change {{"fraction": 0.1..0.35, "source": "ro_gh_plus", "powder_scoops": 1.0}}
- vacuum_substrate {{"intensity": 0.1..1.0}}
- rinse_filter {{"id": "filter"}}
- dose_minerals {{"scoops": 0.25..1.0}}
- dose_ial {{"leaves": 0.25..1.0}}
- test_kit {{"channel": "ph"|"nh3"|"nitrite"|"nitrate"|"gh"|"kh"}}
- buy_instrument {{"kind": "ph_meter"|"do_meter"|"thermometer", "tier": "cheap"|"standard"|"lab"}}
- buy_catalog {{"item": <any item in scorecard.catalog>, "quantity": 1..20}}
- sell_batch {{}}
- check_next {{"after_days": 1..90}}, {{"after_hours": 6}}, or {{"until": "next_day"}}

Prefer gathering missing evidence before risky interventions. Avoid buying equipment unless expected value exceeds its cost. When the tank is stable, use check_next with after_days to advance to biologically meaningful results; do not poll every six hours. Use 1-3 days when monitoring a recent concern and 30-90 days when waiting for breeding or juvenile growth. Do not jump beyond the episode horizon. Return a short reason, action kind, and params_json. params_json must be a string containing the JSON object of action parameters."""

    decision = model_decision(prompt, SCHEMA, "shrimp_farm_action", budget, reasoning_effort)
    raw_params = decision.pop("params_json")
    try:
        decision["params"] = json.loads(raw_params)
    except (json.JSONDecodeError, TypeError):
        decision["params"] = {}
        decision["format_error"] = f"invalid params_json: {raw_params!r}"
    if decision.get("kind") not in ALLOWED_KINDS:
        raise ValueError(f"policy selected unsupported action: {decision}")
    return decision


def react_loop(env, setup, budget, max_turns, max_days, reasoning_effort):
    """Measured episode: setup once, then up to max_turns observe → decide → act."""
    obs = env.reset()
    history = []
    for turn in range(1, max_turns + 1):
        limits = decision_limits(public_observation(obs), budget, turns_used=len(history), max_turns=max_turns, max_days=max_days)
        decision = policy_decision(public_observation(obs), history, budget, limits, reasoning_effort)
        if decision["kind"] == "check_next":
            obs = env.check_next(decision["params"])
        else:
            obs = env.step(ReefAction(kind=decision["kind"], params=decision["params"]))
        history.append({
            "turn": turn,
            "action": decision,
            "observation": public_observation(obs),
            "reward": env.last_reward,
            "done": env.last_done,
        })
        if env.last_done:
            break
    return history