🎮 Grid World Reinforcement Learning

Interactive AI Training Platform - Watch agents learn to navigate complex environments!

Built with Value Iteration & Policy Iteration Algorithms

⚙️ Environment Configuration

3 10
-50 50
0.8 0.99
🧮 Learning Algorithm

Choose the RL algorithm to train the agent

🎯 Quick Actions

📚 How It Works

The Environment:

  • ★ Terminal States: Episode ends (Green +100, Red -50)
  • ◆ Center Reward: Configurable reward cell
  • Step Cost: -1 per move (encourages efficiency)
  • Stochastic Movement: 70% intended, 15% each side

The Algorithms:

  1. Value Iteration: Computes optimal values for all states
  2. Policy Iteration: Alternates evaluation and improvement

Tips:

  • High γ (0.99): Patient, plans ahead
  • Low γ (0.85): Impatient, wants quick rewards
  • Negative center reward: Agent avoids it
  • Positive center reward: Agent seeks it out

📊 Training & Results

Examples
🔲 Grid Size 💎 Center Reward ⏱️ Discount Factor (γ) 🧮 Learning Algorithm