Strongly Condemn This Misleading Astra/JEV Minecraft Demonstration
I am writing this issue because after actually reproducing and inspecting this project, I find the way this demonstration is presented deeply misleading.
This is not a complaint about using a fixed seed or a scripted environment. Those are reasonable engineering choices.
The problem is that the presentation strongly emphasizes Astra + JEV as if they are the reason this Minecraft run works, while the source code reveals that a very large part of the actual solution is already hard-coded into the harness.
I spent several hours testing this myself, and the results made the distinction even more obvious.
1. The Minecraft problem has already been heavily solved before the models act
The repository contains a fixed seed and a pre-surveyed route in:
optimization/nether/config.json
The file contains the exact coordinates for:
- the village
- the supply chests
- preparation beds
- the Nether route
- the Nether exit
- the active End portal
The README itself describes the final route as a fixed surveyed route.
This means the system is not discovering the solution from an unknown Minecraft world.
The environment has already been converted into a highly constrained, pre-mapped problem.
Again, fixed benchmarks are fine.
What is not fine is using the impressive-looking number of model decisions as if that number demonstrates an equivalent level of autonomous problem solving.
2. JEV is not freely controlling Minecraft
The actual architecture is much more restrictive than the presentation initially suggests.
nether-agent.mjs constructs a list of hard-coded candidate actions.
Then models.mjs::decide() sends those candidates to JEV and asks JEV to choose one.
So the real control loop is effectively:
game state -> handcrafted candidate generator -> JEV chooses one candidate -> handcrafted executor
The candidate generator already defines things such as:
- loot this specific chest
- collect this specific bed
- follow this surveyed Nether waypoint
- place this specific bridge block
- approach this specific portal
- enter this specific portal
JEV does not invent these actions.
It chooses among actions that the harness has already decided are legal and useful.
And this creates an even more fundamental issue:
When there is only one candidate, a “JEV decision” is not meaningfully a decision.
The model is effectively being asked to select:
[1]
or something equivalent to:
[1,1]
Calling every one of these a model decision and then emphasizing the total number of decisions gives a very misleading impression of the amount of agency the model actually has.
3. I tested the obvious ablation instead of merely arguing about it
I did not rely solely on the README.
I inspected the source and then modified the system myself.
I tested variants including:
- the original Astra + JEV configuration
- removing Astra
- replacing JEV with a random-number-based selector
The system continued to exhibit the same basic outcome classes:
- successful completion
- death
- getting stuck
In other words, replacing the supposed intelligent decision-makers did not eliminate the observed behavior.
That does not mathematically prove that Astra and JEV contribute exactly zero value.
It does demonstrate something much more important:
The current demonstration does not establish that Astra and JEV are responsible for the demonstrated capability.
The distinction matters.
4. The fixed route and handcrafted control logic are doing an enormous amount of work
The repository contains extensive non-LLM logic for:
- route progression
- pathfinding
- portal construction
- portal entry
- navigation
- inventory handling
- safety handling
- failure cooldowns
- combat behavior
- bed attacks
- dragon-state observation
- action filtering
The models therefore operate inside an already heavily engineered control system.
That is not inherently bad.
But then the experiment should be described accordingly.
A more accurate characterization would be:
A handcrafted Minecraft speedrun harness with LLM-assisted high-level planning and bounded action selection.
That is very different from the impression created by presenting the result as an autonomous Astra/JEV Minecraft agent completing the game.
5. “131 JEV decisions” is not an adequate measure of intelligence
The README highlights:
131 JEV decisions and 35 Astra calls
But the number of calls tells us almost nothing about causal contribution.
Suppose 80 of those decisions contain only one legal candidate.
Then “80 JEV decisions” does not mean that JEV solved 80 difficult Minecraft problems.
It means that JEV was queried 80 times inside a constrained interface.
The correct question is not:
How many times did the model get called?
The correct question is:
What changed when the model was replaced by something simpler?
That is exactly why ablation testing exists.
6. There is a very simple experiment that would settle this
Run the same benchmark under identical conditions with:
- Astra + JEV
- Astra removed
- deterministic action selection
- random action selection
Repeat each condition enough times to report meaningful statistics.
For example:
| Controller |
Runs |
Completion |
Death |
Stuck |
Median time |
| Astra + JEV |
N |
? |
? |
? |
? |
| No Astra |
N |
? |
? |
? |
? |
| Deterministic |
N |
? |
? |
? |
? |
| Random |
N |
? |
? |
? |
? |
If Astra/JEV actually provide substantial value, the difference should show up in the distribution.
If they do not, that should also be visible.
That would be a meaningful experiment.
7. This is why I strongly object to the current presentation
My objection is not that the code uses handcrafted routes.
My objection is that the engineering work that actually makes this particular run possible is mixed together with the model calls and then presented in a way that makes the model contribution look much larger than what the implementation demonstrates.
I consider that presentation seriously misleading.
A model being present in a control loop does not demonstrate that the model caused the result.
A model being queried 131 times does not demonstrate 131 meaningful decisions.
And a successful run on a fixed, pre-surveyed seed does not demonstrate general Minecraft autonomy.
I would strongly encourage the author to publish the ablation results rather than relying on the raw number of Astra/JEV calls as evidence of capability.
Until then, I do not think the current demonstration supports the level of autonomy that its presentation suggests.
Strongly Condemn This Misleading Astra/JEV Minecraft Demonstration
I am writing this issue because after actually reproducing and inspecting this project, I find the way this demonstration is presented deeply misleading.
This is not a complaint about using a fixed seed or a scripted environment. Those are reasonable engineering choices.
The problem is that the presentation strongly emphasizes Astra + JEV as if they are the reason this Minecraft run works, while the source code reveals that a very large part of the actual solution is already hard-coded into the harness.
I spent several hours testing this myself, and the results made the distinction even more obvious.
1. The Minecraft problem has already been heavily solved before the models act
The repository contains a fixed seed and a pre-surveyed route in:
optimization/nether/config.jsonThe file contains the exact coordinates for:
The README itself describes the final route as a fixed surveyed route.
This means the system is not discovering the solution from an unknown Minecraft world.
The environment has already been converted into a highly constrained, pre-mapped problem.
Again, fixed benchmarks are fine.
What is not fine is using the impressive-looking number of model decisions as if that number demonstrates an equivalent level of autonomous problem solving.
2. JEV is not freely controlling Minecraft
The actual architecture is much more restrictive than the presentation initially suggests.
nether-agent.mjsconstructs a list of hard-coded candidate actions.Then
models.mjs::decide()sends those candidates to JEV and asks JEV to choose one.So the real control loop is effectively:
game state -> handcrafted candidate generator -> JEV chooses one candidate -> handcrafted executorThe candidate generator already defines things such as:
JEV does not invent these actions.
It chooses among actions that the harness has already decided are legal and useful.
And this creates an even more fundamental issue:
When there is only one candidate, a “JEV decision” is not meaningfully a decision.
The model is effectively being asked to select:
[1]or something equivalent to:
[1,1]Calling every one of these a model decision and then emphasizing the total number of decisions gives a very misleading impression of the amount of agency the model actually has.
3. I tested the obvious ablation instead of merely arguing about it
I did not rely solely on the README.
I inspected the source and then modified the system myself.
I tested variants including:
The system continued to exhibit the same basic outcome classes:
In other words, replacing the supposed intelligent decision-makers did not eliminate the observed behavior.
That does not mathematically prove that Astra and JEV contribute exactly zero value.
It does demonstrate something much more important:
The current demonstration does not establish that Astra and JEV are responsible for the demonstrated capability.
The distinction matters.
4. The fixed route and handcrafted control logic are doing an enormous amount of work
The repository contains extensive non-LLM logic for:
The models therefore operate inside an already heavily engineered control system.
That is not inherently bad.
But then the experiment should be described accordingly.
A more accurate characterization would be:
That is very different from the impression created by presenting the result as an autonomous Astra/JEV Minecraft agent completing the game.
5. “131 JEV decisions” is not an adequate measure of intelligence
The README highlights:
But the number of calls tells us almost nothing about causal contribution.
Suppose 80 of those decisions contain only one legal candidate.
Then “80 JEV decisions” does not mean that JEV solved 80 difficult Minecraft problems.
It means that JEV was queried 80 times inside a constrained interface.
The correct question is not:
The correct question is:
That is exactly why ablation testing exists.
6. There is a very simple experiment that would settle this
Run the same benchmark under identical conditions with:
Repeat each condition enough times to report meaningful statistics.
For example:
If Astra/JEV actually provide substantial value, the difference should show up in the distribution.
If they do not, that should also be visible.
That would be a meaningful experiment.
7. This is why I strongly object to the current presentation
My objection is not that the code uses handcrafted routes.
My objection is that the engineering work that actually makes this particular run possible is mixed together with the model calls and then presented in a way that makes the model contribution look much larger than what the implementation demonstrates.
I consider that presentation seriously misleading.
A model being present in a control loop does not demonstrate that the model caused the result.
A model being queried 131 times does not demonstrate 131 meaningful decisions.
And a successful run on a fixed, pre-surveyed seed does not demonstrate general Minecraft autonomy.
I would strongly encourage the author to publish the ablation results rather than relying on the raw number of Astra/JEV calls as evidence of capability.
Until then, I do not think the current demonstration supports the level of autonomy that its presentation suggests.