Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “segmented regression model”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

NuGraph2: A Graph Neural Network for Neutrino Event Reconstruction

Neutrino experiments are set to probe some of the most important open questions in physics, from CP violation and the nature of dark matter. The technology of choice for many of these experiments is the liquid argon time projection chamber (LArTPC). In current LArTPC experiments, reconstruction performance often represents a limiting factor for the sensitivity. New developments are therefore needed to unlock the full potential of LArTPC experiments. NuGraph2 is a state of the art Graph Neural Network for reconstruction of data in LArTPC experiments. NuGraph2 utilizes a heterogeneous graph structure, with separate subgraphs of 2D nodes (hits in each plane) connected across planes via 3D nodes (space points). The model provides a consistent description of the neutrino interaction across all planes. NuGraph2 is a multi-purpose network, with a common message-passing attention engine connected to multiple decoders with different classification or regression tasks. These include the classification of detector hits according to the particle type that produced them (semantic segmentation) and the separation of hits from the neutrino interaction from hits due to noise or cosmic-ray background. Additional decoders are being developed, performing tasks such as the regression of the neutrino interaction vertex position. Performance results will be presented based on publicly available samples from MicroBooNE. These include both physics performance metrics, achieving 95% accuracy for semantic segmentation and 98% classification of neutrino hits, as well as computational metrics for training and for inference on CPU or GPU. The status of the NuGraph integration in the LArSoft software framework will be presented, as well as initial studies about model interpretability and injection of domain knowledge.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

NuGraph2: A Graph Neural Network for Neutrino Event Reconstruction

Neutrino experiments are set to probe some of the most important open questions in physics, from CP violation and the nature of dark matter. The technology of choice for many of these experiments is the liquid argon time projection chamber (LArTPC). In current LArTPC experiments, reconstruction performance often represents a limiting factor for the sensitivity. New developments are therefore needed to unlock the full potential of LArTPC experiments. NuGraph2 is a state of the art Graph Neural Network for reconstruction of data in LArTPC experiments. NuGraph2 utilizes a heterogeneous graph structure, with separate subgraphs of 2D nodes (hits in each plane) connected across planes via 3D nodes (space points). The model provides a consistent description of the neutrino interaction across all planes. NuGraph2 is a multi-purpose network, with a common message-passing attention engine connected to multiple decoders with different classification or regression tasks. These include the classification of detector hits according to the particle type that produced them (semantic segmentation) and the separation of hits from the neutrino interaction from hits due to noise or cosmic-ray background. Additional decoders are being developed, performing tasks such as the regression of the neutrino interaction vertex position. Performance results will be presented based on publicly available samples from MicroBooNE. These include both physics performance metrics, achieving 95% accuracy for semantic segmentation and 98% classification of neutrino hits, as well as computational metrics for training and for inference on CPU or GPU. The status of the NuGraph integration in the LArSoft software framework will be presented, as well as initial studies about model interpretability and injection of domain knowledge.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

NuGraph2: A Graph Neural Network for Neutrino Event Reconstruction

Neutrino experiments are set to probe some of the most important open questions in physics, from CP violation and the nature of dark matter. The technology of choice for many of these experiments is the liquid argon time projection chamber (LArTPC). In current LArTPC experiments, reconstruction performance often represents a limiting factor for the sensitivity. New developments are therefore needed to unlock the full potential of LArTPC experiments. NuGraph2 is a state of the art Graph Neural Network for reconstruction of data in LArTPC experiments [https://arxiv.org/abs/2403.11872]. NuGraph2 utilizes a heterogeneous graph structure, with separate subgraphs of 2D nodes (hits in each plane) connected across planes via 3D nodes (space points). The model provides a consistent description of the neutrino interaction across all planes. NuGraph2 is a multi-purpose network, with a common message-passing attention engine connected to multiple decoders with different classification or regression tasks. These include the classification of detector hits according to the particle type that produced them (semantic segmentation) and the separation of hits from the neutrino interaction from hits due to noise or cosmic-ray background. Additional decoders are being developed, performing tasks such as the regression of the neutrino interaction vertex position. Performance results will be presented based on publicly available samples from MicroBooNE. These include both physics performance metrics, achieving 95% accuracy for semantic segmentation and 98% classification of neutrino hits, as well as computational metrics for training and for inference on CPU or GPU. The status of the NuGraph integration in the LArSoft software framework will be presented, as well as initial studies about model interpretability and injection of domain knowledge.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Decayheatml

This code is designed to predict and analyze the decay heat generated in molten salt reactors (MSRs) using a hybrid approach that combines machine learning and segmented polynomial fitting. The accurate prediction of decay heat is essential for reactor safety and the optimization of spent fuel storage. The code operates through several key components: 1) Data Architecture: It incorporates a modular data architecture that handles various MSR-specific operational parameters such as power density, humidity content, and air ingress. These parameters are sampled using Sobol sequences to ensure comprehensive coverage of operational uncertainties. 2) Machine Learning Framework: The code employs a diverse set of machine learning models, including polynomial regression, decision trees, random forests, gradient boosting, support vector regression, k-nearest neighbors, multi-layer perceptrons, and symbolic regression. These models are trained to predict decay heat over a wide temporal range, from immediate shutdown up to 10,000 years. 3) Region-Optimized Training: The temporal domain is divided into multiple regions, each modeled separately to capture distinct decay heat characteristics across different time scales. This approach significantly improves the accuracy and interpretability of predictions. 4) Segmented Polynomial Interpretation (SPI): The SPI method translates machine learning predictions into piecewise polynomial equations. These equations are physically interpretable and can be directly integrated into existing engineering workflows and safety analyses. 5) Front-End Interfaces: The code includes both a Jupyter notebook interface for research development and a Streamlit web application for operational deployment. These interfaces allow users to interactively explore decay heat predictions, adjust operational parameters, and visualize results in real-time. 6) Applications: The framework supports various applications, including safety system validation and spent fuel container optimization. It enables real-time evaluation of worst-case decay heat scenarios, informing the design of passive safety systems and optimizing container designs for long-term storage. Overall, this code provides a robust, accurate, and user-friendly tool for predicting decay heat in MSRs, enhancing reactor safety, and optimizing spent fuel management.

Retamales, Mauricio Eduardo Tano [Idaho National L↗

Combined Data and Deep Learning Model Uncertainties: An Application to the Measurement of Solid Fuel Regression Rate

In complex physical process characterization, such as the measurement of the regression rate for solid hybrid rocket fuels, where both the observation data and the model used have uncertainties originating from multiple sources, combining these in a systematic way for quantities of interest (QoI) remains a challenge. In this paper, we present a forward propagation uncertainty quantification (UQ) process to produce a probabilistic distribution for the observed regression rate r. We characterized two input data uncertainty sources from the experiment (the distortion from the camera U c and the non-zero-angle fuel placement U Y ), the prediction and model form uncertainty from the deep neural network (U m ), as well as the variability from the manually segmented images used for training it (U s ). Here, we conducted seven case studies on combinations of these uncertainty sources with the model form uncertainty. The main contribution of this paper is the investigation and inclusion of the experimental image data uncertainties involved, and how to include them in a workflow when the QoI is the result of multiple sequential processes.

42 ENGINEERING↗

Generating Exploration Mission-3 Trajectories to a 9:2 NRHO using Machine Learning

The purpose of this thesis is to design a machine learning algorithm platform that provides expanded knowledge of mission availability through a launch season by improving trajectory resolution and introducing launch mission forecasting. The specific scenario addressed in this paper is one in which data is provided for four deterministic translational maneuvers through a mission to a Near Rectilinear Halo Orbit (NRHO) with a 9:2 synodic frequency. Current launch availability knowledge under NASA's Orion Orbit Performance Team is established by altering optimization variables associated to given reference launch epochs. This current method can bean abstract task and relies on an orbit analyst to structure a mission based on an established mission design methodology associated to the performance of Orion and NASA's Space Launch System. Introducing a machine learning algorithm trained to construct mission scenarios within the feasible range of known trajectories reduces the required interaction of the orbit analyst by removing the needed step of optimizing the orbit to fit an expected translational response required of the spacecraft. In this study, k-Nearest Neighbor and Bayesian Linear Regression successfully predicted classical orbital elements for the launch windows observed. However both algorithms had limitations due to their approaches to model fitting. Training machine learning algorithms off of classical orbital elements introduced a repetitive approach to reconstructing mission segments for different arrival opportunities through the launch window and can prove to be a viable method of launch window scan generation for future missions.

Guzman, Esteban↗

Data-Efficient Methods for Determining Flory–Huggins χ Parameters in Multicomponent Polymer Formulations

Polymer formulations are essential in diverse applications including personal care products, coatings, paints, adhesives, and plastic materials. Designing these formulations requires navigating large, complex design spaces, where phase and self-assembly behavior critically impact performance. The Flory–Huggins χ parameter, which quantifies segmental miscibility, is widely used to parametrize the excess free energy of mixing in formulation models. In this work, we introduce two data-efficient, top-down methods for estimating χ parameters using the Random Phase Approximation (RPA): (i) Boundary Nonlinear Regression (Boundary-NLR), which fits theoretical spinodal boundaries to experimental phase boundaries, and (ii) Surrogate Model Inverse Parameter Estimation (SMIPE), which uses a Gaussian Process Classifier to fit sparse phase maps via a surrogate model. Both methods allow rapid parametrization of polymer field-theoretic models without the need for additional experiments. We evaluate these approaches on data sets involving polymer–solvent–nonsolvent ternary mixtures and block copolymer–solvent systems, demonstrating their robustness to experimental noise and their relevance for real-world formulation design.

copolymers↗

Collective Risk Ranking of Highway Segments on the Basis of Severity-Weighted Crash Rates

This study is intended to focus on the major factors affecting traffic crash rates and severity levels, in addition to identifying crash-prone locations (i.e., black spots) based on the two indicators. The available crash data for different road segments used for the analysis were obtained from the Washington state database provided by the Highway Safety Information System (HSIS) for the years 2006 to 2011. A Random Forest (RF) classifier was used to predict the outcome level of crash severity, while crash rates were predicted by applying RF regressor. Certain features were selected for each model besides the abstraction of new features to check if there are unobserved correlations affecting the independent variables, such as accounting for the number and weight of crashes within 1 km2 area by implementing the Getis-Ord Gi∗ index. Moreover, to calculate the collective risk (CR) score, crash rates were adjusted to incorporate crash severity weights (cost per severity type) and regression-to-the-mean (RTM) bias via Empirical Bayes (EB) method. Finally, segments were ranked according to their CR score.

Li, Dawei↗

Estimating Volume, Biomass, and Carbon in Hedmark County, Norway Using a Profiling LiDAR

A profiling airborne LiDAR is used to estimate the forest resources of Hedmark County, Norway, a 27390 square kilometer area in southeastern Norway on the Swedish border. One hundred five profiling flight lines totaling 9166 km were flown over the entire county; east-west. The lines, spaced 3 km apart north-south, duplicate the systematic pattern of the Norwegian Forest Inventory (NFI) ground plot arrangement, enabling the profiler to transit 1290 circular, 250 square meter fixed-area NFI ground plots while collecting the systematic LiDAR sample. Seven hundred sixty-three plots of the 1290 plots were overflown within 17.8 m of plot center. Laser measurements of canopy height and crown density are extracted along fixed-length, 17.8 m segments closest to the center of the ground plot and related to basal area, timber volume and above- and belowground dry biomass. Linear, nonstratified equations that estimate ground-measured total aboveground dry biomass report an R(sup 2) = 0.63, with an regression RMSE = 35.2 t/ha. Nonstratified model results for the other biomass components, volume, and basal area are similar, with R(sup 2) values for all models ranging from 0.58 (belowground biomass, RMSE = 8.6 t/ha) to 0.63. Consistently, the most useful single profiling LiDAR variable is quadratic mean canopy height, h (sup bar)(sub qa). Two-variable models typically include h (sup bar)(sub qa) or mean canopy height, h(sup bar)(sub a), with a canopy density or a canopy height standard deviation measure. Stratification by productivity class did not improve the nonstratified models, nor did stratification by pine/spruce/hardwood. County-wide profiling LiDAR estimates are reported, by land cover type, and compared to NFI estimates.

Nelson, Ross↗

Design and Development of a Model to Simulate 0-G Treadmill Running Using the European Space Agency's Subject Loading System

Develop a model that simulates a human running in 0 G using the European Space Agency s (ESA) Subject Loading System (SLS). The model provides ground reaction forces (GRF) based on speed and pull-down forces (PDF). DESIGN The theoretical basis for the Running Model was based on a simple spring-mass model. The dynamic properties of the spring-mass model express theoretical vertical GRF (GRFv) and shear GRF in the posterior-anterior direction (GRFsh) during running gait. ADAMs VIEW software was used to build the model, which has a pelvis, thigh segment, shank segment, and a spring foot (see Figure 1).the model s movement simulates the joint kinematics of a human running at Earth gravity with the aim of generating GRF data. DEVELOPMENT & VERIFICATION ESA provided parabolic flight data of subjects running while using the SLS, for further characterization of the model s GRF. Peak GRF data were fit to a linear regression line dependent on PDF and speed. Interpolation and extrapolation of the regression equation provided a theoretical data matrix, which is used to drive the model s motion equations. Verification of the model was conducted by running the model at 4 different speeds, with each speed accounting for 3 different PDF. The model s GRF data fell within a 1-standard-deviation boundary derived from the empirical ESA data. CONCLUSION The Running Model aids in conducting various simulations (potential scenarios include a fatigued runner or a powerful runner generating high loads at a fast cadence) to determine limitations for the T2 vibration isolation system (VIS) aboard the International Space Station. This model can predict how running with the ESA SLS affects the T2 VIS and may be used for other exercise analyses in the future.

Caldwell, E. C.↗

Space shuttle propulsion parameter estimation using optional estimation techniques

A regression analyses on tabular aerodynamic data provided. A representative aerodynamic model for coefficient estimation. It also reduced the storage requirements for the "normal' model used to check out the estimation algorithms. The results of the regression analyses are presented. The computer routines for the filter portion of the estimation algorithm and the :"bringing-up' of the SRB predictive program on the computer was developed. For the filter program, approximately 54 routines were developed. The routines were highly subsegmented to facilitate overlaying program segments within the partitioned storage space on the computer.

Source record↗

A Centralized AI Lakehouse Framework for Brain Tumor MRI Classification and Segmentation, University KPI Forecasting, and Water Potability Prediction

In many university and healthcare projects, models are built for very different data types such as tables, institutional time series, and medical images, but they are deployed as separate applications. In this work, that separation made testing and maintenance difficult because each module had its own pipeline and runtime requirements. This paper presents an integrated AI lakehouse-style implementation that runs three model pipelines inside one containerized backend. For medical imaging, we used MRI datasets from IEEE DataPort: a four-class classification set with 7012 images (5708 train/1304 test) and a segmentation set with 3063 image–mask pairs. The classification model (ResNet50 transfer learning) is evaluated using a proper train–validation–test protocol across multiple splits (80/10/10, 70/10/20, 60/10/30, and 10/30/60), achieving a test accuracy of 99.00% under the standard 80/10/10 split. Additionally, a patient-level evaluation is conducted using an external glioma dataset to provide a more realistic assessment without data leakage. The segmentation model (DeepLabV3-ResNet50) achieved 83.09% validation mIoU and 88.79% Dice score. For university KPI forecasting, we used annual IPEDS and NSF HERD data from 2010 to 2023 for three universities (BSU, EOU, and UAB). To examine the effect of preprocessing on forecasting performance, two case studies are conducted. In the first case, linear interpolation is applied to generate semester-level data. In the second case, the original annual data is used directly without interpolation. Random Forest regression and ARIMA models are evaluated using MAE, RMSE, MAPE, and R 2 . The results showed that interpolation improved apparent forecasting performance due to smoothing, while evaluation on the original annual data provided a more realistic assessment of model behavior. To further validate the framework on a larger dataset, an additional case study is conducted using a student dropout dataset. For water potability, we trained and compared multiple tabular classifiers on a large dataset (1,048,575 samples). A Random Forest model (100 trees, max depth 10) achieved 85.86% test accuracy and high recall for unsafe samples (0.8447). All modules are served via FastAPI and deployed together using Docker, with workflow automation routing requests to the correct endpoint. System-level benchmarking indicates that the backend maintains stable throughput and latency under concurrent requests.

97 MATHEMATICS AND COMPUTING↗

The Viability of Trajectory Analysis for Diagnosing Dynamical and Chemical Influences on Ozone Concentrations in the UTLS

The viability of trajectory analysis for diagnosing the interplay between chemistry and dynamics is investigated by comparing ozone mixing ratios modelled using air-parcel pathways to values observed along flight tracks during ATTREX (Airborne Tropical TRopopause EXperiment). Trajectories are initiated at the locations of ozone observations and tracked backward in time to their sources at termini of backward trajectories. The modelled values of ozone utilize 3-dimensional analysis fields from WACCM (Whole Atmosphere Community Climate Model) (a chemical-climate model with dynamical fields nudged towards MERRA (Modern-Era Retrospective Analysis and Research Applications) reanalysis) and ERA-interim (product of ECMWF - the European Centre for Medium-Range Weather Forecasts) to determine source mixing ratios with chemical production and loss terms derived from the ozone chemistry used in WACCM. A statistical base of modelled ozone is constructed with 6 trajectory platforms (adiabatic, diabatic, and kinematic forced by ERA-interim and MERRA), two chemical models (WACCM chemistry and no chemistry), and 4 trajectory lengths (5, 10, 20, and 30 days). Linear regression is employed to separate systematic errors from random errors and to characterize the impact of source mixing ratios, path length, vertical motion, and chemistry on modelled ozone errors. Errors in the analysis ozone fields are large, if not dominant, contributors to model error. Random errors are particularly large for point-by-point comparisons, however averaging over 800 km (75 minutes) flight segments substantially reduces random error and exposes systematic errors. Of the two analysis ozone data sets, WACCM, which incorporates detailed chemistry, provides the smaller systematic errors while ERA-interim, which has crude chemistry but assimilates observational data, has the smaller random errors. Of the different trajectory platforms, adiabatic calculations produce the smaller random errors (irrespective of the use of chemistry) but both vertical motion and chemistry are required to optimally reduce systematic errors. These results suggest that meaningful analysis of dynamical and chemical interactions that control ozone mixing ratios are viable on spatial scales larger than a few reanalysis grid spaces, that errors in the analyzed ozone data sets are large but not prohibitively so, and that vertical velocities and heating rates from reanalysis data, while problematic, contain useful information [on the ozone concentrations in the UTLS (Upper Troposphere/Lower Stratosphere)].

ozone↗

Contributions of Astronauts Aerobic Exercise Intensity and Time on Change in VO2peak during Spaceflight

There is considerable variability among astronauts with respect to changes in maximal aerobic capacity (VO2peak) during International Space Station (ISS) missions, ranging from a 5% increase to 30% decline. Individual differences may be due to in-flight aerobic exercise time and intensity. PURPOSE: To evaluate the effects of in-flight aerobic exercise time and intensity on change in VO2peak during ISS missions. METHODS: Astronauts (N=11) performed peak cycle tests approx 60 days before flight (L-60), on flight day (FD) approx 14, and every approx 30 days thereafter. Metabolic gas analysis and heart rate (HR) were measured continuously during the test using the portable pulmonary function system. HR and duration of each in-flight cycle ergometer and treadmill (TM) session were recorded and averaged in time segments corresponding to each peak test. Mixed effects linear regression with exercise mode (TM or cycle) as a categorical variable was used to assess the contributions of exercise intensity (%time >70% peak HR or %time >90% peak HR) and time (min/wk), adjusted for body weight, on %change in VO2peak during the mission, and incorporating the repeated-measures experimental design. RESULTS: 110 observations were included in the model (4-6 peak cycle tests per astronaut, 2 exercise devices). VO2peak was reduced from preflight throughout the mission (FD14: 13+/-13% and FD 105: 8+/-10%). Exercise intensity (%peak HR: FD14=66+/-14; FD105=75+/-8) and time (min/wk: FD14=82+/-46; FD105=158+/-40) increased during flight. The models showed main effects for exercise time and intensity with no interactions between time, intensity, and device (70% peak HR: time [z-score=2.39; P=0.017], intensity [z-score=3.51; P=0.000]; 90% peak HR: time [zscore= 3.31; P=0.001], intensity [z-score=2.24; P=0.025]). CONCLUSION: Exercise time and intensity independently contribute to %change in VO2peak during ISS missions, indicating that there are minimal values for exercise time and intensity required to maintain VO2peak. As the FD105 average exercise intensity and time did not prevent a decline in VO2peak from preflight, astronauts' exercise prescriptions should target at least 160 min of weekly aerobic exercise at an average above 75% peak HR with increased time at intensities above 90% of peak HR starting early in the mission.

Downs, Meghan E.↗

Machine learning enhanced characterization and optimization of photonic cured MAPbI 3 for efficient perovskite solar cells

Photonic curing (PC) can facilitate high-speed perovskite solar cell (PSC) manufacturing because it uses high-intensity light pulses to crystallize perovskite films in milliseconds. However, optimizing PC conditions is challenging due to its many variables, and using power conversion efficiency (PCE) as the optimization metric is both time-consuming and labor-intensive. This work presents a machine learning (ML) approach to optimize PC conditions for fabricating methylammonium lead iodide (MAPbI 3 ) films by quantitatively comparing their ultraviolet-visible (UV-vis) absorbance spectra to thermal annealed (TA) films using four similarity metrics. We perform Bayesian optimization coupled with Gaussian process regression (BO-GP) to minimize the similarity metrics. Refining PC conditions using active learning based on BO-GP models, we achieve a PC MAPbI3 film with an absorbance spectrum closely matching a TA reference film, which is further verified by its crystalline and morphological properties. Thus, we demonstrate that the UV-vis absorption spectrum can accurately proxy film quality. Additionally, we use an AI-based segmentation model for a more efficient grain size analysis. However, when we use the optimized PC condition to fabricate PSCs, we find that interaction between MAPbI 3 and the hole transport layer (HTL) during PC critically degrades the PSC performance. By adding a buffer layer between the HTL and MAPbI 3 , the optimized PC PSCs produce a champion PCE of 11.8%, comparable to the TA reference of 11.7%. Using UV-vis similarity metrics instead of device PCE as the objective in our BO-GP method accelerates the optimization of PC processing conditions for MAPbI 3 films.

14 SOLAR ENERGY↗

Estimation of EOP From VLBI: Direct Approach

The currently adopted strategy of Earth Orientation Parameters (EOP) estimation from Very Long Baseline Interferometry (VLBI) is to estimate six parameters: Universal Time 1 (UT1), UT1 rate, pole positions, and nutation offsets for each 24-hour session independently. Then the resulting time series of raw EOP are filtered and a regression analysis is performed to obtain nutation coefficients, polhode of the pole, and other physical parameters. Thus, the latter parameters are obtained indirectly in two stages. An alternative approach of direct estimation of the final EOP is presented. Pole coordinates and UT1 are considered as a sum of three components: the low-period component that is modeled by a cubic spline, the harmonic component that includes forced nutation, precession and sub-daily variations of EOP, and the stochastic component that is modeled by a linear spline with segment length 1-2 hours. All parameters are obtained in a single LSQ solution using all available data.

POLAR MOTION↗

Three-dimensional segmentation of luminal and adventitial borders in serial intravascular ultrasound images

Intravascular ultrasound (IVUS) provides exact anatomy of arteries, allowing accurate quantitative analysis. Automated segmentation of IVUS images is a prerequisite for routine quantitative analyses. We present a new three-dimensional (3D) segmentation technique, called active surface segmentation, which detects luminal and adventitial borders in IVUS pullback examinations of coronary arteries. The technique was validated against expert tracings by computing correlation coefficients (range 0.83-0.97) and William's index values (range 0.37-0.66). The technique was statistically accurate, robust to image artifacts, and capable of segmenting a large number of images rapidly. Active surface segmentation enabled geometrically accurate 3D reconstruction and visualization of coronary arteries and volumetric measurements.

NASA Discipline Cardiopulmonary↗

Machine learning approaches for structural and thermodynamic properties of a Lennard-Jones fluid

Predicting the functional properties of many molecular systems relies on understanding how atomistic interactions give rise to macroscale observables. However, current attempts to develop predictive models for the structural and thermodynamic properties of condensed-phase systems often rely on extensive parameter fitting to empirically selected functional forms whose effectiveness is limited to a narrow range of physical conditions. Here, we illustrate how these traditional fitting paradigms can be superseded using machine learning. Specifically, we use the results of molecular dynamics simulations to train machine learning protocols that are able to produce the radial distribution function, pressure, and internal energy of a Lennard-Jones fluid with increased accuracy in comparison to previous theoretical methods. The radial distribution function is determined using a variant of the segmented linear regression with the multivariate function decomposition approach developed by Craven et al. [J. Phys. Chem. Lett. 11, 4372 (2020)]. The pressure and internal energy are determined using expressions containing the learned radial distribution function and also a kernel ridge regression process that is trained directly on thermodynamic properties measured in simulation. The presented results suggest that the structural and thermodynamic properties of fluids may be determined more accurately through machine learning than through human-guided functional forms.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗