To do
PredPreyGrass observation-space improvements
Investigate whether agents need derived internal-state and ecological-context features in addition to the current raw observations.
Current setup:
local grid + own/other energy → movement
Agents currently have to learn by trial and error that:
energy < starvation threshold → danger
energy > reproduction threshold → reproduction opportunity
nearby allies + combined energy → possible group kill
nearby prey → chase value
grass density direction → foraging opportunity
enemy pressure direction → escape/flee pressure
Possible improvement:
local grid + raw energy + derived drives/context → movement
Start with minimal biologically plausible derived features:
hunger_pressure
reproductive_readiness
danger_pressure
social_pressure / isolation
Avoid adding overly engineered features too early, such as:
can_kill_this_prey = true
best_escape_direction = north
best_grass_direction = east
Reason:
Raw observation:
more open-ended, less biased, harder to learn
Derived internal drives:
more sample-efficient, biologically plausible, still not too hand-designed
Detailed affordances:
faster learning perhaps, but more assumptions injected
Suggested experiment order:
1. Baseline:
local grid + energy → movement
2. Drive-conditioned version:
local grid + energy + hunger/reproduction/danger/social pressure → movement
3. Optional later:
add ecological affordance features such as kill feasibility, grass gradient, escape availability
4. Compare:
episode_len_mean
all-types-survive-to-horizon rate
predator/prey extinction timing
birth/death rates
group-hunt events
population stability
lineage persistence
Main principle:
Keep the action space movement-only.
Make behavior richer by improving what the agent observes,
not by adding explicit actions like eat, reproduce, or attack.
Yes — this can be very useful for your Darwinian/Baldwinian two-timescale approach.
The key idea is:
Fast timescale / Nurture:
agents learn movement behavior during their lifetime
Slow timescale / Nature:
evolution selects the inherited structures that make learning easier or more adaptive
Derived observation features such as:
hunger_pressure
reproductive_readiness
danger_pressure
social_pressure
ally_support
local_food_density
can become part of the evolvable interface between biology and learning.
Why this fits the Baldwinian idea
The Baldwin effect is about how learning within a lifetime can influence evolutionary selection. In computational terms: agents learn during evaluation, but what gets selected is not necessarily the learned behavior itself; rather, selection favors inherited traits that make useful learning easier. Reviews of the Baldwin effect describe this as adaptive learning affecting the direction or speed of evolution, although the exact effect can be accelerating or constraining depending on the setup. (PMC)
For PredPreyGrass, that maps very nicely:
Inherited/evolved:
observation structure
drive thresholds
hunger sensitivity
danger sensitivity
social sensitivity
reproduction threshold interpretation
initial policy biases / architecture
Learned during lifetime:
actual movement policy
when to approach prey
when to avoid danger
when to stay near allies
when to forage
So your agents would not inherit “hunt this way” directly. Instead, they may inherit a better learning scaffold.
Example:
Lineage A:
sees only raw energy
Lineage B:
sees hunger_pressure and danger_pressure
Lineage C:
sees hunger_pressure, danger_pressure, and social_pressure
Then PPO learning happens inside each lineage. If Lineage C learns stable pack hunting and population survival faster, evolution may select Lineage C. That is very Baldwinian.
The clean two-timescale architecture
You can frame it like this:
Outer loop: Darwinian / Baldwinian selection
mutate inherited traits
train each variant for some number of PPO iterations
evaluate survival/reproduction/cooperation/population stability
keep successful variants
Inner loop: reinforcement learning
given inherited observation structure and drive parameters
learn movement-only behavior
So the inner loop remains:
observation → movement
but the outer loop evolves what kind of observation/drive system the agent is born with.
What exactly can evolve?
You do not have to evolve only neural network weights. You can evolve the biological interpretation layer.
For example:
hunger_threshold
danger_radius
social_radius
grass_density_radius
reproduction_pressure_scaling
enemy_pressure_scaling
ally_pressure_scaling
whether a drive channel is enabled or disabled
observation range
movement speed variant
energy cost
initial policy seed / architecture
This is very interesting for your project because it lets evolution discover which “motivational systems” are useful.
For example:
Variant 1:
hunger-sensitive, weakly social
Variant 2:
danger-sensitive, strongly social
Variant 3:
reproduction-sensitive, low danger sensitivity
Variant 4:
raw-energy only, no derived drives
Then you can ask:
Which inherited drive structure produces better lifetime learning?
Which drive structure supports stable predator-prey coexistence?
Which drive structure supports group hunting?
Which drive structure supports longer lineage persistence?
That is exactly in line with your Nature/Nurture framing.
Darwinian vs Baldwinian in your model
You can separate them like this:
Pure Darwinian version
Evolution selects inherited traits based on fitness.
Little or no within-lifetime learning.
Example:
mutate drive parameters
evaluate behavior
select variants with more reproduction/survival
This tests what can be selected directly.
Baldwinian version
Each variant learns during its lifetime.
Selection uses post-learning fitness.
But learned weights are not directly inherited.
Example:
variant is born with hunger/danger/social drive settings
variant trains with PPO for N iterations
evaluate learned behavior
select variants whose inherited setup made learning successful
offspring inherit drive settings, not the learned final policy
This tests whether evolution can favor learnability.
Lamarckian version, optional comparison
Learned policy weights are inherited directly.
This is less biologically realistic, but useful as a computational benchmark.
Why this may help your project
It gives your two-timescale process a much cleaner experimental object.
Instead of saying:
evolution selects policies
you can say:
evolution selects motivational/perceptual scaffolds
learning fills in the movement behavior
That is a stronger biological model.
It also connects to hierarchical RL. Temporal abstraction in RL is often used to represent behavior at multiple levels, where higher-level structures organize lower-level actions over time. Sutton, Precup, and Singh’s options framework is a classic example of this kind of temporal abstraction. (www-anw.cs.umass.edu)
For PredPreyGrass, the hierarchy could eventually become:
Evolution:
selects drive systems and learning biases
High-level learned mode:
forage / flee / group / reproduce-oriented movement
Low-level learned policy:
concrete movement direction
But you do not need to start with full hierarchy. The first step is simpler:
evolve observation/drive parameters
train movement policy with PPO
select variants by ecological fitness
Best experiment design
I would add this to your roadmap:
Experiment: Baldwinian observation-drive evolution
Goal:
Test whether evolved internal-drive features improve lifetime learning
and population-level stability.
Inner loop:
PPO learns movement-only behavior using:
local grid
raw energy
derived drive channels
Outer loop:
mutate/select drive parameters:
hunger scaling
danger scaling
social-radius
reproduction-readiness scaling
enabled/disabled drive channels
Fitness:
all-types-survive-to-horizon rate
reproduction success
lineage persistence
predator/prey coexistence
group-hunt events
population stability
The most important conceptual payoff is this:
Raw observation asks:
Can learning discover useful behavior from scratch?
Derived drives ask:
Can evolution produce perceptual/motivational systems
that make useful learning easier?
That is very close to your Darwinian/Baldwinian approach.
Model Hunter Gathererers
- What are good determinants of "Camp"
Hub formatioen
- as a bridge between hunter gateherers and settlers
- What are good determinats of "Hub" (for more permanent settlement)
- Water/river
- Protection
- (Proximity access) to Leaderchip (Marbella: first the elite tourism then mass toerism; Hapton court)
- Scaleability/self fulfilling
- Decisison to Fight-or-Flight
eating
=- More varied than setllers
- Scavanging
- Nuts
- Deer
- Large Deer (stronger than humans)
sheletering
household formation
movement
bands
- size (150)
- social structure
- specializasing
steps
-every step 1 month, to acurately enough simulate the seasons.
Mental accounting
Implement a range of in-debtness, to model the concept of friend/informal business. For instance if defected x-times in a row / or if the accumulated investement is 'fair'.
ESS : with only thieves in the world there is nothing to steal
if defectors get punished with a certain probability; how does that reduce crime? [Rachel: "commiting crime is inversely related to chance of being caught/punished"]
make a vizualization with manually inserting a strange startegy into a basin/gird
- To (dis) proof ESS
Nature v Nurture definitions
-
what is nurture ? Pure "self nurtured" or "man made nurtured" or "nature nurtured? If someone is born near the equator in Africa; is that nurture? Is the behavior of ancestors nurture or nature? Is a physical inheritance nurture or nature?
-
"Humans seem to have evolved capacities for learning reciprocity, but the actual reciprocal rules are built through development, attachment, repeated interaction, and culture. A newborn does not “reciprocate” in the game-theory sense. A baby mainly receives care. But babies are already equipped for social responsiveness: attention to faces, voices, emotional expressions, turn-taking rhythms, comfort, attachment, and sensitivity to contingent responses. Those are evolved foundations."
| Nowak mechanism | “Nature” side | “Nurture” side | Human version |
|---|---|---|---|
| Kin selection | Strong. Selection favors helping genetic relatives because they share genes. | Moderate. Humans learn who counts as “family,” and culture can expand or weaken kin duties. | “I help my child, sibling, cousin, clan.” |
| Direct reciprocity | Strong-medium. Evolution favors capacities for memory, trust, gratitude, resentment, partner recognition. | Strong. Individuals learn who helps, who cheats, who can be trusted. | “You helped me before, so I help you now.” |
| Indirect reciprocity | Medium. Evolution favors reputation tracking, moral emotions, concern for social evaluation. | Very strong. Reputation depends on language, gossip, norms, morality, cultural rules. | “You helped others, so I trust/help you.” |
| Network reciprocity | Medium. Evolution can favor clustering, bonding, local loyalty, partner choice. | Strong. Human networks are shaped by family, school, work, religion, online groups, institutions. | “People in my circle help each other.” |
| Group selection / multilevel selection | Medium-strong. Groups with more internal cooperation may outcompete less cooperative groups. | Very strong in humans. Group identity, norms, punishment, rituals, laws, ideology, and institutions are culturally transmitted. | “We cooperate because we are part of this group.” |
- Kin selection is the most “nature-heavy.” Indirect reciprocity and group-level cooperation are the most “nurture/culture-heavy.” Direct reciprocity sits in the middle. But none of them is purely nature or purely nurture.
A more useful division is this:
| Level | What evolution supplies | What learning/culture supplies |
|---|---|---|
| Basic social architecture | attachment, social emotions, memory, recognition, fairness sensitivity, punishment motives | — |
| Development | readiness to learn social rules | who helped me, who cheated, whom to trust |
| Culture | capacity for norm learning | actual rules: family duty, fairness, debt, gratitude, honor, punishment |
| Institutions | capacity for group living | law, religion, markets, reputation systems, contracts |
So in humans, evolutionary selection does not simply produce fixed cooperation rules like:
“Always cooperate.”
That would be too rigid and easily exploited.
Instead, selection favors learning systems that allow humans to become cooperative under the right social conditions: cooperate with kin, reciprocate with reliable partners, care about reputation, cluster with cooperators, follow group norms, and punish or avoid exploiters.
Why do humans cooperate?
The surface answers are many:
- survival,
- reciprocity,
- empathy,
- norms,
- reputation,
- laws,
- morality.
These are important mechanisms, but they can be reduced to a smaller set of structural reasons.
Evolutionary Stable Strategy (ESS)
- "With only thieves (in the world) there's nothing to steal"
Having options makes people happy
- Changing (options) seasons makes people happie than peoplewith fixed climate?? This implies relation between equator distance and happiness.
Interdependence of outcomes
Cooperation becomes rational when payoffs are coupled and agents cannot optimize fully on their own.
Examples include:
- shared resources,
- division of labor,
- ecological feedback loops,
- public goods,
- tasks that exceed solo capacity.
Temporal extension
Cooperation becomes more likely when interactions repeat over time. Short-term sacrifice can produce long-term gain through:
- reciprocity,
- trust,
- reputation,
- learning,
- cultural transmission.
Internalization of group structure
Humans often carry social regulation inside the individual through:
- empathy,
- guilt,
- shame,
- norms,
- identity.
This helps explain why cooperation can persist even when direct monitoring or immediate reward is weak.
One-sentence synthesis
Cooperation emerges when independent optimization breaks down, the future matters, and social coordination becomes internalized.
Why this matters here
For this site, the central question is not whether cooperation exists, but how it fits into a broader theory of human behavior and under what minimal conditions it emerges:
- through learning within a lifetime,
- through selection across generations,
- and through the interaction between those two timescales.
That is why cooperation sits near the center of the project. It is not the whole of human behavior, but it is one of the clearest cases in which behavior cannot be understood at the level of the isolated individual alone.
Brainstorming
- Use Leary's Rose in Learned Cooperation?
- "The Inevitablity of Selfishness"
- "Cooperation is not trivial. Competition is intuitivly more sensible due to the inevitability of selfishness"
[X]Cooperation by bundled forces of predators
- Predators can eat if the predators enters the Moore neighborhood of a Prey
- It can only eat if it has a higher energy level than the Prey
- If more Predators ar in the Moore neighborhood, it can eat the Prey only if their cumulative energy is greater or equal than the Prey. If so: the divide the energy proportionally to their own energy.
- "When you can't do it alone you must do it together"
Similarties "Nature" and "Nurture"
- maybe not so different
- The natural selection of life time learning
- diminshing returns on the happy behaviors (learning is also open-ended like evolution is)
- reward system is adaptivre like evolution
Differences
- "nature" is very binary: survival/reproduction, "Nurture" is more continuous and less fatal.
repo tit-for-tat
- iterative tit-for-tat
- always end defecting if finite ending
- MARL training with random ending periods (max_steps), might result in cooperating
[ ] Direct reciprocity without coordination under necessity
Goal:
- Remove "we must cooperate or we cannot kill the prey" completely.
- Study whether predators learn to help because help is returned later, not because a kill is impossible alone.
Recommended environment concept:
- Start from a rabbits-only or
shared_preystyle environment, notmammoths. - Every prey is individually catchable by one predator.
- Reproduction remains the only learning reward, so the setup stays aligned with the rest of PredPreyGrass.
Cooperative act:
- After a successful solo kill, the capturing predator can choose
share_food = 0/1. - If
share_food = 1and another predator is within Moore neighborhood: - A fixed fraction of prey energy is transferred to one nearby predator.
- The sharer keeps the remainder and is immediately worse off than under selfish consumption.
- Sharing is therefore voluntary and immediately costly.
Alternative cooperative act:
assist_hunt = 0/1for a nearby predator that is chasing prey.- Assistance lowers the target predator's hunting cost or raises its capture chance.
- Assistance is never required for capture, only beneficial.
Direct reciprocity mechanism:
- Each predator keeps private memory of specific partners, not public reputation.
- Example memory variable:
trust[i][j]= how much predatoriexpects predatorjto return favors. - Increase
trust[i][j]whenjshared with or assistedi. - Decrease
trust[i][j]whenjrefused to share or help in a relevant opportunity. - Let trust slowly decay back toward neutral so reciprocity must be maintained.
Observation / state:
- Standard spatial observation stays intact.
- Add one extra private observation signal for predators only:
- At nearby predator positions, encode focal-agent trust toward that predator.
- Or provide a compact summary such as nearest-partner trust / mean nearby trust.
- Do not expose a public reputation score; otherwise the mechanism shifts toward indirect reciprocity.
Why this is no longer necessity:
- A predator can always eat alone.
- Cooperation now means giving up immediate energy for another predator.
- The only reason to do this is expectation of future return through repeated interaction.
Core experimental conditions:
- Baseline selfish condition: no memory, no partner-specific trust signal.
- Direct reciprocity condition: private partner memory enabled.
- Identity-shuffle ablation: same reciprocity logic, but predator identities are randomly remapped each episode.
- Optional indirect reciprocity comparison: public reputation signal instead of private pairwise memory.
Ecological settings that make direct reciprocity testable:
- Spawn offspring near parents so the same predators meet repeatedly.
- Keep movement costs and energy decay moderate so repeated interaction matters.
- Keep prey abundant enough that sharing is feasible, but not so abundant that social help is irrelevant.
- Keep lifetimes long enough for remembered favors to be returned.
What should emerge if direct reciprocity is real:
- Predators share or assist reliable partners more than unreliable partners.
- Predators reduce helping after a partner failed to reciprocate.
- Cooperation is stronger with partner memory than without it.
- Cooperation collapses or weakens strongly when identities are shuffled.
Minimal metrics:
P(share | partner shared with me before)P(share | partner did not share with me before)P(assist | partner assisted me before)- Mean energy transferred per dyad over time
- Share/assist rate for familiar partners versus unfamiliar partners
- Change in helping probability after partner defection
- Reproduction rate under baseline vs reciprocity vs identity-shuffle
Interpretation:
- If helping rises only when partner-specific memory is available, then cooperation is no longer explained by immediate ecological necessity.
- It is explained by expected future return from repeated interaction: direct reciprocity.
[ ] Mixed -Stah Hunt
- Have to types of Prey: Mammoths and Deer
- Experiment for coevolution
- https://chatgpt.com/share/694e5758-e21c-8008-87d9-1c01dc66cf1b
- https://en.wikipedia.org/wiki/Stag_hunt
Macro-level energy
- Add it to the file which is already in place: energy_by_type.json (created by evaluate_......_debug.py)
- Substract cumulative decay energy Predator and Prey per step (homeostatic energy)
- Add cumulative photosynthesis energy from grass
layered cooperationn in SocialBehavior
- Marl Book example
- Display Maslow's Pyramid and describe project from bottom to top:
- First layer: PredatorPreyGrass project. Typical for the first layer, physical need (eeating), survival an reproduction.
- Second layer: social needs. Need to corporate
Dynamic training
-
Create training algorithm of competing policies and select 'winner' after each iteration/a number of iterations. Competing policies have different environment configs. Goal: optimize environment parameters more efficiently and automatically at run time rather than manually after full (10 hour) experiments. Determine success:
- fitness metrics
- ability to co-adapt
-
curriculum reward tuning
Examples to try out
-
Meta learning example RLlib ("learning-to-learn"): https://github.com/ray-project/ray/blob/master/rllib/examples/algorithms/maml_lr_supervised_learning.py
-
curriculum: https://github.com/ray-project/ray/blob/master/rllib/examples/curriculum/curriculum_learning.py
-
curiosity: https://github.com/ray-project/ray/tree/master/rllib/examples/curiosity
-
Explore examples: https://github.com/flairox/jaxmarl?tab=readme-ov-file
Environment enhancements
-
Male & Female reproduction instead of asexual reproduction
-
Build wall or move wall
-
Adding water/rivers
Experiments
-
Tuning hyperparameters and env parameters simultaneously (see chat)
-
max_steps_per_episode: For policy learning performance: 500–2000 steps per episode is a common sweet spot in multi-agent RL — long enough for interactions to unfold, short enough for PPO to assign credit.
For open-ended co-evolution (your case): you might intentionally want longer episodes (e.g. 2000–5000) so emergent dynamics have time to play out, even if training is slower.
A good trick is to curriculum the horizon:
Start short (e.g. 500–1000) → agents learn basic survival.
Gradually increase (e.g. +500 every N iterations) → expose them to longer ecological timescales.
“works-in-practice” plan for your PredPreyGrass run, plus what to tweak as you lengthen episodes.
Recommended episode horizon + hyperparameters (curriculum)
Start shorter for stability/throughput, then stretch to let eco-dynamics (booms, busts, Red-Queen) unfold.
Phase A (bootstrap)
max_steps = 1_000gamma = 0.995(effective credit horizon ≈ 1/(1−γ) ≈ 200 steps)lambda_ (GAE) = 0.95–0.97
Phase B (mid)
max_steps = 2_000–3_000gamma = 0.997–0.998(horizon ≈ 333–500)lambda_ = 0.96–0.97
Phase C (long-term dynamics)
max_steps = 4_000–5_000gamma = 0.998–0.999(horizon ≈ 500–1 000)lambda_ = 0.97
Why that mapping? PPO’s useful credit horizon is ~1/(1−γ). As you increase
max_steps, you raiseγso actions can “see” far enough ahead without making variance explode.Batch/throughput knobs to adjust as episodes get longer
Keep ~4–10 episodes per PPO iteration so you still get decent reset diversity:
- train_batch_size: roughly
episodes_per_iter × max_steps. Example: atmax_steps=1_000, use8_000–16_000. When you move tomax_steps=3_000, bump toward24_000–48_000. - rollout_fragment_length: increase with horizon so GAE has longer contiguous fragments (e.g., 200 → 400 → 800).
- num_envs_per_env_runner: raise a bit as episodes lengthen to maintain sampler throughput.
- KL/clip: leave defaults unless you see instability; longer horizons often benefit from slightly smaller learning rate rather than big clip/kl changes.
When to stop stretching episodes
- If
timing/iter_minutesballoons or TensorBoard curves update too slowly, hold the currentmax_stepsfor a while. - If you see extinction before the cap, longer episodes won’t help—tune ecology (e.g., energy gains/losses) instead.
Make available the BHP archive in a repository
LT-goal acquire more wealth as a population
- Energy as a proxy of wealth
- Only the top 10% of energy reproduces?
- Escaping the Malthusian trap
Integrate Dynamic Field Theory
- Wrapper around brain
- Visualize first!!!
Posting on the linkedin The Behavior Patterns Project?
The Malthusian Trap in a Predator–Prey Co-Evolutionary System
- Limit population size of predators or prey; is that beneficial compared to unbounded reproduction?
- [2.5]: experiment_1 / experiment are powerfull examples of the Malthusian trap. Record this on site. Create a "Malthusian Trap"
Pranjal 2-12-2025
- Communication: leave ant trace (ant colony/lenia), also keep previous state in Observation?
- www.talkRL.com
- Reshape Field of Vision for Predators; only into the direction of moiving? In that way Prey can hide more easily from Predators?
- Is the existence of a prolonged episode betweeen predators and prey not an emergence of cooperation?
Research shortlist: evolution + birth/death + MARL
-
Malthusian Reinforcement Learning (Leibo et al., 2018/2019): population pressure and ecology-linked MARL adaptation.
https://arxiv.org/abs/1812.07019
https://www.ifaamas.org/Proceedings/aamas2019/pdfs/p1099.pdf -
Neural MMO (Suarez et al., 2019; Neural MMO 2.0, 2021): persistent many-agent worlds with spawn/death and resource pressure.
https://arxiv.org/abs/1903.00784
https://arxiv.org/abs/2110.07594 -
Evolutionary Population Curriculum (2020): evolutionary selection over policy populations in large-scale MARL.
https://arxiv.org/abs/2003.10423 -
Evolutionary MARL in Group Social Dilemmas (Chaos, 2025): evolutionary pressure on RL traits in social dilemmas.
https://pubmed.ncbi.nlm.nih.gov/39937196/ -
Iterated + Evolutionary Games with MARL (Nature Communications, 2025): MARL-discovered strategies tested in evolving populations.
https://www.nature.com/articles/s41467-025-67178-6 -
Neural Population Learning beyond Symmetric Zero-Sum Games (AAMAS 2024): population-level selection/equilibrium in general-sum MARL.
https://deepmind.google/research/publications/24820/ -
Comenius and Curriculum Learning
Comenius argued that teaching should proceed from the easy to the difficult, so that new knowledge builds on what has already been learned. This principle closely resembles curriculum learning in reinforcement learning, where an agent first trains on simpler tasks before progressing to more complex ones.
For PredPreyGrass, this could mean starting with easy survival conditions and gradually introducing scarcity, predators, competition, cooperation, and co-evolution. An adaptive curriculum may be especially useful, because the difficulty can change according to the agents’ current performance rather than following a fixed sequence. This connects Comenius’ educational principle with modern ideas in automatic curriculum learning and open-ended learning.
- make a to do mindmap?