Formulation scientists are accustomed to working with incomplete information.
A new adhesive, coating, polymer compound or personal-care formulation may contain five, ten or twenty adjustable variables, while laboratory time allows only a small fraction of the possible combinations to be tested. Development therefore depends on experimental strategy as much as chemical knowledge.
Design of Experiments has helped R&D teams manage this problem for decades. Properly designed factorial, response-surface and mixture experiments can reveal important main effects, interactions and useful operating regions with far fewer runs than uncontrolled trial and error.
Bayesian optimization does not make Design of Experiments obsolete. It addresses a somewhat different problem: what should be tested when the experimental plan is allowed to change after every new result?
That distinction is becoming increasingly relevant as chemical R&D moves toward autonomous experimentation.
Traditional Experimental Designs Are Usually Planned Before the Results Arrive
In a conventional Design of Experiments programme, researchers define factors, ranges, responses and an experimental design before the work begins.
The planned experiments are then carried out, the data are analysed and a model is fitted. Researchers may subsequently design another experiment if refinement is needed.
This approach works extremely well when the objective is understanding factor effects, building interpretable models and exploring a reasonably well-defined experimental region.
The limitation appears when experiments are expensive and the design space is very large.
Suppose a formulation team is optimizing an electrically conductive adhesive. The variables include epoxy-to-hardener ratio, silver loading, particle-size distribution, reactive diluent, accelerator level, dispersant concentration and cure temperature. Responses include viscosity, electrical resistance, lap-shear strength and glass-transition temperature.
The team does not necessarily need a complete statistical map of every possible interaction. It may primarily need to identify a formulation meeting several performance limits with as few experiments as possible.
This is where Bayesian optimization becomes attractive.
Bayesian Optimization Learns While the Experimental Campaign Is Running
Bayesian optimization typically uses a surrogate model to estimate how experimental variables relate to a target response.
More importantly, the model also represents uncertainty.
An acquisition function then decides which experiment is most useful to perform next. Sometimes that means testing a formulation predicted to perform very well. At other times it means testing a poorly understood region because reducing uncertainty there may improve the overall model.
After the result is collected, the surrogate model is updated and another experiment is selected.
The experimental sequence therefore becomes adaptive.
A July 14, 2026 paper in npj Computational Materials introduced a unified Bayesian optimization framework intended to accelerate materials discovery, reflecting the continued expansion of these approaches across complex materials spaces.
Recent work is also showing the value of combining Bayesian optimization with scientific knowledge rather than treating the system as an entirely blind search. An ACS study published in July incorporated researchers' prior knowledge directly into the optimization framework to identify useful materials conditions with fewer experiments.
For industrial formulators, that is particularly important because decades of chemistry knowledge already exist.
Formulation Problems Are Especially Suitable for Adaptive Optimization
Many formulation problems share characteristics that make Bayesian optimization useful.
Experiments may be expensive. The design space is multidimensional. Variables interact. The relationship between composition and performance may be nonlinear. Several responses may conflict with one another.
A polyurethane adhesive might need high strength but low modulus. A thermal interface material may need maximum thermal conductivity while remaining dispensable. A coating could require corrosion resistance, flexibility, adhesion and low volatile-organic-compound content simultaneously.
The optimum therefore rarely lies at the maximum of one response.
Bayesian optimization can be configured for multiple objectives or constraints, allowing researchers to explore trade-offs rather than optimizing a single property in isolation.
Recent materials research is increasingly applying this logic to practical development problems. A study published August 28, 2026 used Bayesian optimization to explore process-structure-property relationships in polyethylene composite films, demonstrating how adaptive optimization can guide materials development across coupled variables.
Bayesian Optimization and DoE Should Not Be Treated as Competitors
One of the least useful discussions in AI-enabled R&D is the claim that machine learning will replace Design of Experiments.
Experienced industrial teams are more likely to benefit from combining the two.
Design of Experiments provides structured exploration, interpretable factor effects and statistically disciplined experimental planning. Bayesian optimization provides adaptive selection of future experiments based on what has already been learned.
A practical workflow might begin with a space-filling or statistically designed initial dataset. Bayesian optimization can then use that data to recommend subsequent experiments.
The first stage gives the model broad information about the system. The second stage concentrates experimental resources where improvement is most promising.
This combination can be particularly valuable when datasets are small, which is common in formulation R&D.
Small Data Changes the AI Strategy
Much discussion about artificial intelligence assumes very large datasets.
Formulation scientists often have the opposite problem.
A project may contain 40 reliable historical experiments, not 40 million. Some formulations were tested for viscosity but not adhesion. Raw-material grades changed during the programme. Processing conditions may not have been recorded consistently.
This makes conventional deep-learning approaches difficult.
Bayesian methods are attractive partly because they can operate in relatively small-data environments while representing uncertainty explicitly.
The model does not need to pretend that it knows everything.
When uncertainty is high, the optimization strategy can deliberately choose experiments that provide additional information.
That is a powerful idea for R&D because the best experiment is not always the formulation predicted to have the highest performance. Sometimes the most valuable experiment is the one that reduces uncertainty enough to improve future decisions.
Failed Experiments Still Contain Information
Traditional project records frequently contain a bias toward successful formulations.
Researchers may document the final candidates carefully while failed batches, unstable samples or unprocessable formulations receive minimal attention.
An autonomous optimization system needs those failures.
A formulation that phase separated immediately provides useful information about the feasible design region. A sample that exceeded the maximum viscosity limit tells the model that certain combinations should be avoided. A cure experiment that produced excessive exotherm may define a safety constraint.
Negative results therefore become part of the model's understanding of the design space.
This requires better data discipline because the reason for failure matters. A formulation that failed chemically is different from an experiment invalidated by an instrument malfunction or weighing error.
The transition to AI-guided formulation therefore depends as much on experimental metadata as on algorithms.
Multi-Objective Optimization Is Where Industrial Value Becomes Clearer
Academic optimization problems are often expressed as finding a maximum or minimum response.
Industrial formulation is rarely that simple.
A commercially useful product may need to meet ten specifications simultaneously while remaining cost-effective and manufacturable.
Instead of asking for the formulation with maximum peel strength, the optimization problem might be defined as:
increase peel strength while keeping viscosity below the dispensing limit, maintaining glass-transition temperature above the requirement, reducing raw-material cost and avoiding cure temperatures incompatible with the customer's process.
The output may not be one perfect formulation.
It may be a Pareto front containing several technically attractive compromises.
The formulator then brings business, manufacturing and application knowledge into the final decision.
This is a good example of how artificial intelligence can augment rather than eliminate human judgement.
Formulation Optimization Becomes Even More Powerful When Connected to Automation
Bayesian optimization can be used without laboratory robotics.
Scientists can manually prepare the experiments recommended by the algorithm and return the measured results.
Once preparation and analytics become automated, however, the optimization loop can run more frequently.
A self-driving laboratory combines this adaptive decision layer with automated execution and measurement. The algorithm recommends the next experiment, laboratory equipment executes it, analytical instruments produce feedback and the model updates its recommendation.
Current self-driving laboratory research increasingly describes precisely this closed-loop relationship between AI, automation and experimentation.
The key point for formulation teams is that Bayesian optimization is useful before the laboratory becomes fully autonomous.
It can be one of the first practical steps toward agentic R&D.
The Question Is Not How Many Experiments AI Can Eliminate
Claims that an optimization algorithm can reduce experiments by a fixed percentage should be treated cautiously.
Performance depends heavily on the problem, data quality, variable selection, constraints, starting dataset and underlying response landscape.
A simple formulation may already be efficiently solved using conventional statistical design. A highly nonlinear system with expensive measurements may benefit substantially from adaptive optimization.
The objective should not be to demonstrate that artificial intelligence uses fewer experiments at any cost.
It should be to increase the information gained from each experiment.
That is a much more useful R&D metric.
Bringing Bayesian Optimization Into Chemical R&D
For formulation teams, the practical questions are often more difficult than the mathematical definition of Bayesian optimization.
Which formulation variables should be included? How should categorical raw-material choices be represented? How should hard constraints be handled? How much initial data are required? Which responses should be optimized simultaneously? How should failed experiments be encoded? When should scientists override the recommendation?
These implementation decisions determine whether the system becomes scientifically useful.
The OnlyTRAININGS advanced session Agentic AI in Chemical R&D: Self-Driving Labs & Formulation Optimization examines Bayesian optimization in this wider experimental context.
The training connects active learning and optimization algorithms with formulation design, experimental feedback, AI agents, laboratory automation, data quality and human oversight so that R&D teams can understand how adaptive optimization fits into real chemical development workflows.
Explore the Agentic AI in Chemical R&D training:
https://www.onlytrainings.com/course/agentic-ai-chemical-rd-self-driving-labs-formulation-optimization/
Bayesian optimization, formulation optimization, AI formulation optimization, Design of Experiments, active learning, multi-objective optimization, materials optimization, AI-driven formulation, experimental design, Bayesian optimization for chemical formulation, Bayesian optimization vs Design of Experiments, AI formulation optimization for chemical R&D, machine learning formulation optimization, Bayesian optimization materials discovery
