NCKH – Pilot Scale Reinforcement Learning
Reinforcement Learning (RL) Forecast Dashboard
Policy-based forecaster that evaluates control-style actions from telemetry history
Air Temperature31,03°C
Air Humidity71,65%
Wind Speed13,82 km/h
Rainfall0,26 mm
ConditionsSunny intervals
Auto API
Air Temperature31,03 °C
Air Humidity71,65 %
Wind Speed13,82 km/h
Rainfall0,26 mm
ConditionsSunny intervals
ESP32 + BME280 + BH1750 + Rain gauge
28,22
7,97
6,75
4,11
31,03
71,65
13,82
0,26
Forecast Window:
━━ Observed   ╌╌ Reinforcement Learning (RL)
Water Temperature (°C)
PH Parameter
Dissolved Oxygen (mg/L)
Turbidity (NTU)
Salinity (ppt)
TDS (mg/L)
ORP (mV)
25,1
7,52
7,51
2,56
31,03
71,65
13,82
0,26
Forecast Window:
━━ Observed   ╌╌ Reinforcement Learning (RL)
Water Temperature (°C)
PH Parameter
Dissolved Oxygen (mg/L)
Turbidity (NTU)
Salinity (ppt)
TDS (mg/L)
ORP (mV)
0.126
0.176
0.955
2.04
Training Loss vs Validation Loss MSE per epoch
MAE Convergence by Epoch
R² Score theo Epoch
Predicted vs Actual – Scatter (Temperature °C)
Model Architecture – Reinforcement Learning (RL)
Reinforcement Learning Policy Network
  • State: multi-sensor window (temperature, pH, DO, turbidity, weather)
  • Policy backbone: 3 x Dense(96), ReLU activation
  • Action: next-step parameter adjustment
  • Reward: +stability, +threshold compliance, -overshoot
  • Training style: offline replay buffer + periodic online update
  • Loss: policy + value objectives (actor-critic style)
Control-Oriented Design
  • Adaptive to short-term pond dynamics
  • Balances trend-following and stabilization
  • Penalty for threshold crossing risk
  • Uses weather context as exogenous signal
  • Supports continual policy improvement from telemetry