🎮 Grid World Reinforcement Learning
Interactive AI Training Platform - Watch agents learn to navigate complex environments!
Built with Value Iteration & Policy Iteration Algorithms
⚙️ Environment Configuration
3 10
-50 50
0.8 0.99
🎯 Quick Actions
📚 How It Works
The Environment:
- ★ Terminal States: Episode ends (Green +100, Red -50)
- ◆ Center Reward: Configurable reward cell
- Step Cost: -1 per move (encourages efficiency)
- Stochastic Movement: 70% intended, 15% each side
The Algorithms:
- Value Iteration: Computes optimal values for all states
- Policy Iteration: Alternates evaluation and improvement
Tips:
- High γ (0.99): Patient, plans ahead
- Low γ (0.85): Impatient, wants quick rewards
- Negative center reward: Agent avoids it
- Positive center reward: Agent seeks it out
📊 Training & Results
Compare Value Iteration vs Policy Iteration side-by-side
Examples
| 🔲 Grid Size | 💎 Center Reward | ⏱️ Discount Factor (γ) | 🧮 Learning Algorithm |
|---|
🔬 Experiment Ideas
Risk vs Reward: Try γ=0.99 vs γ=0.85 to see patience vs greed
Avoidance: Set center reward to -30 and watch the agent avoid it
Complexity: Increase grid size to 8x8 for challenging navigation
Comparison: See which algorithm converges faster for different configurations
Built with ❤️ using Reinforcement Learning | Gradio | Python
Deploy on Hugging Face Spaces • Perfect for learning RL concepts • Interactive & Educational