Reagent & Reinforcement-Learning Optimisation

Dose reagents better, recover from upsets faster — learned policies

Application

Real-time supervisory control keeps a flotation circuit optimised moment to moment, but two longer-horizon questions remain: how much reagent should be dosed, and what should the control strategy do when the circuit is recovering from an upset? Both benefit from learning over time rather than reacting in the instant. Minealytics layers two complementary learning capabilities on top of the cell supervisor, supported by AI-assisted tuning of the optimiser itself.

Reagent Management

Machine-learning models provide recommendations for collector and frother dosing — for example in the rougher circuit — based on the observed state of the feed and the froth. These recommendations complement existing expert systems rather than replacing them, giving operators a data-driven view of dosing alongside the rules and heuristics they already trust.

Reinforcement-Learning Policy Learner

The reinforcement-learning policy learner acts as a “slow optimiser”. It studies historical upset and recovery events and learns the best “recovery recipes” — the ideal initial targets and air splits that previously stabilised similar upsets.

When the circuit exits an unstable state, the policy learner prescribes the recipe that worked best in comparable situations, jump-starting the fast optimiser from a good position rather than letting it search from scratch. Because it learns continuously from every new event, it becomes the long-term memory that guides real-time control, steadily improving the way the circuit recovers from disturbances.

AI-Assisted Tuning

A well-behaved optimiser depends on well-chosen weights, learning rates and pause thresholds. Minealytics calibrates these offline so the live system is never the place where parameters are guessed:

  • Parameter calibration: optimiser weights, learning rates and pause thresholds are tuned offline against historical data.
  • Contextual parameter packs: the system selects parameter packs suited to the operating regime — for example by feed-solids band or frother dosage — so the optimiser behaves appropriately as conditions change.
  • Collapse-risk surrogates: short-horizon surrogate models predict froth-collapse risk and reject poor candidate moves before they are ever tried on the plant.

Broader Objectives

Because the optimiser works to an explicit objective function, the goals it pursues can be broadened beyond recovery alone. Reagent cost, total air-capacity limits and downstream thickener constraints can all be incorporated, shifting the aim from purely maximising recovery toward optimising economic profit across the circuit.

Operating mode

Configured per site

The approved operating mode depends on the site, available data, validation results and safety case. A capability may begin as monitoring or advice, then progress to supervised or closed-loop control. Existing PLC, DCS and safety interlocks remain the final authority on what equipment can do.

Read about deployment and assurance