A rigorous, validation-first engineering framework for building trustworthy physics-based simulations and AI-assisted scientific software.
The number of unknowns in simulation—and specifically in plasma simulation—has been growing year over year. Now, AI provides an unprecedented opportunity to tackle this challenge, provided that computational frameworks adopt FAIR data standards to ensure reproducibility and enable physics-informed architectures [1, 2].
It is about tracking the sheer volume of unknowns. While there are standard studies on sensitivity analysis, the challenge goes much further because there are inherent errors in cross-sections and other fundamental inputs. The lack of machine-readable community standards for plasma chemistries hinders the precise reconstruction and verification of published models [3].
The WHY aims to explain how we can tackle these issues head-on. One of our core goals at Deep Why is to perform rigorous simulation engineering while remaining acutely aware of the rapidly shifting landscape—the future of AI is already here, and everything is changing very fast.
Do you want to get ready with us? Join the WHY.
In computational modeling and AI engineering, precise definitions ensure clarity between technical teams, stakeholders, and automated decision tools:
Verification asks: "Was the software implemented correctly?" It ensures equations are solved accurate to numerical tolerances without bugs.
Validation asks: "Does the model represent the real-world system accurately enough for its intended use?" It benchmarks model predictions against physical experiments or literature data.
AI-Powered: Leveraging AI tools to accelerate code generation, automated benchmarking, and parametric exploration.
AI-Agentic Methodology: Autonomous software agents executing modular tasks (e.g. sweep generation, regression runs) under explicit domain constraints.
Integrating fundamental physical laws (conservation of mass, momentum, energy, Maxwell's equations) directly into solver architectures or machine learning models to guarantee physical consistency.
Human-guided domain expert oversight reviewing automated validation reports, boundary conditions, and uncertainty margins to make final engineering decisions.
Our methodology structures scientific software development into a continuous validation pipeline:
Establish quantitative error tolerances, intended physical regimes, and acceptance criteria before writing code.
Construct analytical test cases, exact benchmark solutions, and unit test suites to confirm solver mechanics.
Curate experimental measurements and published literature benchmarks for rigorous empirical comparison.
Build domain models incorporating chemical kinetic schemes, EM field coupling, and transport equations.
Accelerate code modernizations (e.g. MATLAB to Python) and boilerplate generation using specialized AI agent prompts.
Parity tests against reference baselines run as the code changes; full benchmark sweeps are rerun for each campaign.
Quantify input parameter uncertainty and evaluate model stability across operating conditions.
Domain experts evaluate physical plausibility, edge-case behaviors, and trade-offs.
Generate transparent, reproducible audit reports linking validation benchmarks to source code releases.
The nine steps are not run once. The work is organised in campaigns: each campaign is a complete, reproducible run against a fixed validation target, and each one widens what is covered — more of the codebase, more of the physics, more independent solvers. The residuals a campaign leaves behind define the scope of the next one.
Fix the benchmark: the published figures and reference results the code must reproduce.
Run the full parametric study with a recorded configuration, so it can be reproduced.
Compare every observable against the reference, with errors quoted from the output files.
Separate what is measured from what is suspected. Open residuals become the hypotheses of the next campaign.
This campaign approach is not specific to one model: it can be applied to any open-source scientific codebase.
Worked example: pyFrost-GM, a Python reimplementation of the LoKI-GM global model, validated against Dias et al. (2023) and Alves et al. (2026). Two campaigns are complete and a third is planned. The second campaign added an independent Boltzmann solver and found seven integration defects along the way.
We help plasma-technology companies, scientific-software teams, advanced manufacturing companies, and R&D departments develop simulation tools that are testable, traceable, and fit for critical engineering decisions.