Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “regression models”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

The impact of kidney function on Alzheimer’s disease blood biomarkers: implications for predicting amyloid-β positivity

Impaired kidney function has a potential confounding effect on blood biomarker levels, including biomarkers for Alzheimer’s disease (AD). Given the imminent use of certain blood biomarkers in the routine diagnostic work-up of patients with suspected AD, knowledge on the potential impact of comorbidities on the utility of blood biomarkers is important. We aimed to evaluate the association between kidney function, assessed through estimated glomerular filtration rate (eGFR) calculated from plasma creatinine and AD blood biomarkers, as well as their influence over predicting Aβ-positivity. We included 242 participants from the Translational Biomarkers in Aging and Dementia (TRIAD) cohort, comprising cognitively unimpaired individuals (CU; n = 124), mild cognitive impairment (MCI; n = 58), AD dementia (n = 34), and non-AD dementia (n = 26) patients all characterized by [ 18 F] AZD-4694. Plasma samples were analyzed for Aβ42, Aβ40, glial fibrillary acidic protein (GFAP), neurofilament light chain (NfL), tau phosphorylated at threonine 181 (p-tau181), 217 (p-tau217), 231 (p-tau231) and N-terminal containing tau fragments (NTA-tau) using Simoa technology. Kidney function was assessed by eGFR in mL/min/1.73 m 2 , based on plasma creatinine levels, age, and sex. Participants were also stratified according to their eGFR-indexed stages of chronic kidney disease (CKD). We evaluated the association between eGFR and blood biomarker levels with linear models and assessed whether eGFR provided added predictive value to determine Aβ-positivity with logistic regression models. Biomarker concentrations were highest in individuals with CKD stage 3, followed by stages 2 and 1, but differences were only significant for NfL, Aβ42, and Aβ40 (not Aβ42/Aβ40). All investigated biomarkers showed significant associations with eGFR except plasma NTA-tau, with stronger relationships observed for Aβ40 and NfL. However, after adjusting for either age, sex or Aβ-PET SUVr, the association with eGFR was no longer significant for all biomarkers except Aβ40, Aβ42, NfL, and GFAP. When evaluating whether accounting for kidney function could lead to improved prediction of Aβ-positivity, we observed no improvements in model fit (Akaike Information Criterion, AIC) or in discriminative performance (AUC) by adding eGFR to a base model including each plasma biomarker, age, and sex. While covariates like age and sex improved model fit, eGFR contributed minimally, and there were no significant differences in clinical discrimination based on AUC values. We found that kidney function seems to be associated with AD blood biomarker concentrations. However, these associations did not remain significant after adjusting for age and sex, except for Aβ40, Aβ42, NfL, and GFAP. While covariates such as age and sex improved prediction of Aβ-positivity, including eGFR in the models did not lead to improved prediction for any biomarker. Our findings indicate that renal function, within the normal to mild impairment range, does not seem to have a clinically relevant impact when using highly accurate blood biomarkers, such as p-tau217, in a biomarker-supported diagnosis.

60 APPLIED LIFE SCIENCES↗

Accurate prediction of carbon dioxide capture by deep eutectic solvents using quantum chemistry and a neural network

Carbon dioxide (CO 2 ) emissions from fossil fuel combustion are a significant source of greenhouse gas, contributing in a major way to global warming and climate change. Carbon dioxide capture and sequestration is gaining much attention as a potential method for controlling these greenhouse gas emissions. Among the environmentally friendly solvents, deep eutectic solvents (DESs) have demonstrated the potential capability for carbon capture. To establish a theoretical framework for DES activity, thermodynamics modeling and solubility predictions are significant factors to anticipate and understand the system behavior. Here, in this study, we combine the COSMO-RS model with machine learning techniques to predict the solubility of CO 2 in various deep eutectic solvents. A comprehensive data set was established comprising 1973 CO 2 solubility data points in 132 different DESs at a variety of temperatures, pressures, and DES molar ratios. This data set was then utilized for the further verification and development of the COSMO-RS model. The CO 2 solubility (ln(x CO 2 )) in DESs calculated with the COSMO-RS model differs significantly from the experiment with an average absolute relative deviation (AARD) of 23.4%. A multilinear regression model was developed using the COSMO-RS predicted solubility and a temperature-pressure dependent parameter, which improved the AARD to 12%. Finally, a machine learning model using COSMO-RS-derived features was developed based on an artificial neural network algorithm. The results are in excellent agreement with the experimental CO 2 solubilities, with an AARD of only 2.72%. The ML model will be a potentially useful tool for the design and selection of DESs for CO 2 capture and utilization.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Gaussian-process generative model for the QCD equation of state

We develop a generative model for the nuclear matter equation of state at zero net baryon density using the Gaussian process regression method. We impose first-principles theoretical constraints from lattice quantum chromodynamics and hadron resonance gas at high- and low-temperature regions, respectively. By allowing the trained Gaussian process regression model to vary freely near the phase transition region, we generate random smooth crossover equations of state with different speeds of sound that do not rely on specific parametrizations. Here, we explore a collection of experimental observable dependencies on the generated equations of state, which paves the groundwork for future Bayesian inference studies to use experimental measurements from relativistic heavy-ion collisions to constrain the nuclear matter equation of state.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

A Nonlinear Regression Method for Composite Protection Modeling of Induction Motor Loads

Protection equipment are used to prevent damages to induction motor loads by isolating those from the power network in the event of severe faults. Modeling the response of induction motor loads and their protection is vital for power system planning and operation, especially in understanding system's response moments after a fault has occurred. This article proposes an optimization based framework to generate composite protection models for commercial building motor loads. Introducing a mathematical abstraction, the task of finding a suitable (simplified) model of the composite protection scheme is formulated as a nonlinear regression problem. Numerical examples are provided to illustrate the application of the framework.

Kundu, Soumya↗

Sub-pilot-scale Production of High-Value Products from U.S. Coals

Investigators from the University of Utah, University of Wyoming and Marshall University pursued a program to study the conversion of raw coal to high-value products of carbon fiber and silicon carbide. Team members also developed an initial framework for a data portal that can incorporate laboratory data on coal processing and product quality, and also work with tools for machine learning for data analysis, data visualization and economic assessment. Experimental R&D efforts focused on the conversion of raw coal to coal tar and other byproducts, and the resulting tar intermediates were upgraded to form anisotropic and isotropic pitch materials. These pitch materials were produced from coal using both thermal (pyrolysis) and chemical (mild solvolysis liquefaction) decomposition of raw coal. Four different coals were studied: Utah bituminous coal (Sufco), Wyoming PRB coal (Black Thunder), Illinois bituminous coal (Illinois #6), and West Virginia bituminous coal (Flying Eagle). Both metallurgical-grade coking coals and lower-grade steam coals were investigated, and controlled secondary gas-phase reactions were used during a two-stage pyrolysis process to induce cracking and condensation reactions among the pyrolytic tar species. This approach successfully improved the performance of the lower grade coals for yielding pitch materials, with properties more consistent with a commercial-grade pitch that had previously demonstrated success for quality carbon fiber production. The use of waste plastic materials was also studied, to help improve physical and chemical characteristics of the intermediate tars and final pitch product; in particular, for lowering the pitch softening point to an acceptable level for melt spinning carbon fiber. Mild solvolysis liquefaction was also used as a method for producing pitch for carbon fiber production. As expected, significantly higher pitch yields were obtained using this approach, and waste plastic materials were also successfully used to reduce pitch softening point to an acceptable level. The plastic materials were also utilized to create a solvent for the mild solvolysis process, and this plastic-derived solvent was shown to provide results consistent with more expensive commercial chemical solvents, and could thus avoid the need for costly recovery and recycle of a liquefaction solvent. Additional experimental R&D focused on the production of silicon carbide (β-SiC) from the residual char byproduct from pitch production, and also on the production of carbon fiber from the anisotropic pitch. SiC was successfully synthesized using a mixture of residual char and sandstone at a ratio of 1:1. Reaction temperature and residence time were optimized and yielded a product purity of 81%. For carbon fiber production, the most successful pitch samples were obtained from the mild solvolysis liquefaction approach, combined with the use of a plastic (HDPE)-derived solvent. Fiber properties improved over time as laboratory fiber production methodologies improved, and final yields of carbon fiber were obtained with a diameter of 12.14 ± 1.10 um, Modulus of 173.73 ± 15.25 GPa, and Tensile Strength of 1.04 ± 0.10 GPa. A proof-of-concept Modern Community Research Data Portal (MCRDP) was developed and deployed for coal and coal-derived pitch characterization, with the full support of (i) remote web-based access, (ii) distributed analysis, (iii) interactive visualization and exploration, (iv) shared and long-term data access, (v) advanced query capabilities and (vi) real-time collaboration. The Coal to Products Data Portal “coaltoproducts.org” provides researchers with space to store and share data within a project, tools for analyzing and understanding data for scientific investigation, and the ability to publish data to the broader community for reproducibility. The portal leverages the Material Commons 2.0 (MC) platform developed by the Center for PRedictive Integrated Structural Materials Science (PRISMS) of the University of Michigan, to achieve long-term longevity of data collections and, more importantly, collaborative science. A number of data visualization tools were also assessed and implemented for interrogating the experimental and modeling data. The machine learning portion of this project analyzed datasets from two different coal conversion processes performed on a diverse set of coal samples from both the coal pyrolysis experiments and the solvent liquefaction experiments. The work was initiated by exploring standard regression models on the pyrolysis data, aiming to understand the impact of sample characteristics and processing conditions on key product metrics. Over the course of the project, the focus expanded to include a variety of machine learning tools, delving into both supervised and unsupervised learning methods. Models tested on the pyrolysis data included linear, ridge, lasso, elastic-net, Gaussian process, random forest regression, and AutoSklearn, and the approach was continually refined to enhance predictive accuracy and model interpretability. Similar techniques were applied to the liquefaction data with an additional focus on feature engineering. Along with mesophase content, additional outputs of interest were the pitch yield, softening point, and QI content. Insights derived from these analyses are crucial in determining the factors influencing the quality and yield of coal-derived products. As the work progressed, the research evolved from foundational model comparisons to analyses of random forests, decision paths, and feature importance scores. A thorough market analysis was performed to examine the prospects of coal-based carbon fibers. The best opportunities for coal come from its lower and more stable price relative to petroleum, particularly for subbituminous coals, which is the primary advantage that a coal refinery may have over a petroleum refinery. Before a commercial CTP production facility can be modeled, however, several things need to be understood regarding the nature of the would-be coal refinery. These include the technology to be deployed, the size of facility, the volume(s) of co-product(s), and the waste and emissions profile of the plant. The volume of co-products and waste may be substantial and will require separate market analysis to ensure viability. In the near-term, the importance of coal tar pitch, in the form of carbon pitch, to the aluminum and steel industries is likely to overshadow the alternative use of this material as an input for carbon fiber. The importance of steel and aluminum in building materials, and the need for carbon materials in their manufacturing, will ensure that demand for these products remains for the long run. In addition, carbon fiber may also be the best substitute for steel and aluminum well into the future. While society will eventually be able to shift production of much of its electricity needs to renewables, it will not be able to shift away from fossil fuels for production of high-strength construction and vehicular materials. Demand for carbon fiber is expected to increase quickly, but the volume of carbon fiber and the amount of coal that would be needed to produce even a sizeable share of this market may still be relatively small compared to current coal production. Thus, other coal-based products like graphene, graphite, carbon foams, resins, and carbon-based building products will play important roles in sustaining coal production as coal-fired power generation continues to decline.

01 COAL, LIGNITE, AND PEAT↗

PyTREES

PyTREES (Python tool for Training/Testing Robust Explainable Ensembles on Spectra) is software that implements a data-driven approach to predicting the amount of specific oxides present in materials samples of laser-induced breakdown spectroscopy (LIBS); such as from the ChemCam instrument suite onboard the NASA Curiosity rover. PyTREES is designed to input LIBS data in the format provided by the ChemCam team [1]. PyTREES then applies appropriate pre-processing to this data [2], and implements several regression methods for predicting oxides from spectra. The regression methods include: ensemble methods (random forest, extra trees, and gradient boosting regression) and blended submodels using the “double blending” technique. PyTREES additionally implements methods for quantifying the importance of features in regression model: (1) mean decrease in impurity (MDI) and (2) permutation importance to investigate the wavelengths used by the regression methods. [1] Gasda et al. (2021). Spectrochim Acta B, 181, 106223. [2] Clegg et al. (2017). Spectrochim Acta B , 129, 64–85.

Oyen, Diane↗

A Swing of the Pendulum: The Chemodynamics of the Local Stellar Halo Indicate Contributions from Several Radial Merger Events

We find that the chemical abundances and dynamics of APOGEE and GALAH stars in the local stellar halo are inconsistent with a scenario in which the inner halo is primarily composed of debris from a single massive, ancient merger event, as has been proposed to explain the Gaia-Enceladus/Gaia Sausage (GSE) structure. The data contain trends of chemical composition with energy that are opposite to expectations for a single massive, ancient merger event, and multiple chemical evolution paths with distinct dynamics are present. We use a Bayesian Gaussian mixture model regression algorithm to characterize the local stellar halo, and find that the data are fit best by a model with four components. We interpret these components as the Virgo Radial Merger (VRM), Cronus, Nereus, and Thamnos; however, Nereus and Thamnos likely represent more than one accretion event because the chemical abundance distributions of their member stars contain many peaks. Although the Cronus and Thamnos components have different dynamics, their chemical abundances suggest they may be related. We show that the distinct low- and high-α halo populations from Nissen & Schuster are explained by VRM and Cronus stars, as well as some in situ stars. Because the local stellar halo contains multiple substructures, different popular methods of selecting GSE stars will actually select different mixtures of these substructures, which may change the apparent chemodynamic properties of the selected stars. We also find that the Splash stars in the Solar region are shifted to higher v $\phi$ and slightly lower [Fe/H] than previously reported.

79 ASTRONOMY AND ASTROPHYSICS↗

Robust Adaptive Decentralized Dynamic State Estimation with Unknown Control Inputs using Field PMU Measurements

This paper proposes a robust adaptive decentralized dynamic state estimation method for power system with unknown inputs of the highly detailed synchronous machine model. The temporal and spatial correlations among the unknown inputs are used to derive a vector auto-regressive model. The latter is further integrated together with state transition and measurement models for joint state and unknown inputs estimation. Thanks to the consideration of implicit cross-correlations between the states and the unknown inputs, only generator terminal voltage and current phasors are needed. Test results on the US WECC system using the field PMU measurements show that the proposed method is able to track both the system dynamic states and unknown controller inputs. These information could significantly benefit the validation and calibration of generator controller parameters.

Zhao, Junbo↗

A cross-dimensional analysis of data-driven short-term load forecasting methods with large-scale smart meter data

Electricity load forecasting is essential to utility operation and power grid stability. A wide spectrum of data-driven methods, ranging from linear regression models to more recent deep learning models have been adopted to forecast electric load over the years. However, there still lacks a holistic evaluation of the applicability of conventional statistical and machine learning based algorithms with respect to different temporal and spatial scopes, computational requirements, and sensitivity of model-tuning. Enabled by a large-scale electricity load profile dataset of over 40,000 residential customers in a utility region, we conducted a cross-dimensional analysis of data-driven load forecasting methods. Three regression-based and seven deep learning algorithms with different model configurations were evaluated in terms of their overall and peak load prediction accuracy, and training burdens, across spatial aggregation levels ranging from the transformer, feeder, substation, to neighborhood. We found, first, the load forecasting accuracy is constrained by a predictability boundary, influenced by the forecasting horizon and spatial aggregation level. Specifically, RandomForest, XGBoost, TFT, TSMixer, and TiDE models achieved less than 10 % prediction error for up to 96-h ahead forecasting for district, substation, and feeder levels, while other models struggle at long-horizon predictions; Second, for winter and summer peak load dates, most models were able to predict the peak demand timing within ± 1 h, but the prediction percentage error varied by models, with TFT and TiDE models being the top performers; Third, models with similar prediction accuracy can differ in training burden by an order of magnitude. Therefore, choosing model configurations that balance prediction performance and computational resource is an important practical consideration for large-scale deployment of the machine learning based load forecasting. The outcome of this study can guide researchers and practitioners to choose the proper load forecasting algorithms based on their problem scope, required accuracy, and available resources. The predictability boundary can serve as a benchmark for electricity load forecasting problems with new algorithms and datasets.

Li, Han↗

Knowledge of lactation amenorrhea method among postpartum women in Ethiopia: a facility-based cross-sectional study

While the importance of knowledge about contraceptives in improving their utilization and thereby reducing the risk of unintended pregnancies is well documented, there are limited studies documented about the Lactational Amenorrhea Method (LAM). Thus, understanding the knowledge of postpartum mothers about LAM is essential for designing tailored interventions. This study assessed the level of knowledge about LAM and its associated factors among postpartum mothers in Ethiopia. A facility-based cross-sectional study was conducted among 3148 randomly selected postpartum participants. The study utilized multistage sampling approach in hospitals located across five regions and one city administration in Ethiopia. Data were collected using face-to-face interviews at discharge. A participant was categorized as having knowledge of LAM if she correctly answered the three LAM criteria: amenorrhea, the first 6 months, and exclusive breast feeding. A binary logistic regression model was used to identify factors associated with knowledge of LAM. Variables with p < 0.25 in the binary logistic regression were included in the multiple logistic regression. Then, associations were described using the adjusted odds ratio (AOR) along with the 95% confidence interval (CI), and statistical significance was declared at p < 0.05. Only four in 10 participants (40.6%; 95% CI 38.9–42.3) had knowledge of LAM. Participants who attended college or above educational level (AOR = 2.1, 95% CI 1.5–2.8), those with parity of two (AOR = 2.3; 95% CI 1.6–3.6) or more than two (AOR = 2.4; 95% CI 1.5–4.0), those who expressed a desire for further fertility (AOR = 1.3; 95% CI 1.1–1.5), individuals who received counselling on LAM (AOR = 3.0; 95% CI 2.6–3.7), and those who gave birth in hospital (AOR = 2.6; 95% CI 1.4–2.6) had higher odds of knowledge about LAM, compared to their counter parts. In contrary, participants resided far away from health facilities had 30% lower odd of knowledge about LAM compared to those resided near the health facilities (AOR = 0.70; 95% CI 0.6–0.8). The proportion of participants who had knowledge of LAM was low. Strengthening counseling about LAM during antenatal care and delivery with due attention to women with limited access to health facilities should be considered for increasing their level of knowledge on LAM.

60 APPLIED LIFE SCIENCES↗

Partner with a Third-Party Delivery Service or Not? A Prediction-and-Decision Tool for Restaurants Facing Takeout Demand Surges During a Pandemic

Amidst the COVID-19 pandemic, restaurants become more reliant on no-contact pick-up or delivery ways for serving customers. As a result, they need to make tactical planning decisions such as whether to partner with online platforms, to form their own delivery team, or both. In this paper, we develop an integrated prediction-decision model to analyze the profit of combining the two approaches and to decide the needed number of drivers under stochastic demand. We first use the susceptible-infected-recovered (SIR) model to forecast future infected cases in a given region and then construct an autoregressive-moving-average (ARMA) regression model to predict food-ordering demand. Using predicted demand samples, we formulate a stochastic integer program to optimize food delivery plans. We conduct numerical studies using COVID-19 data and food-ordering demand data collected from local restaurants in Nuevo Leon, Mexico, from April to October 2020, to show results for helping restaurants build contingency plans under rapid market changes. Our method can be used under unexpected demand surges, various infection/vaccination status, and demand patterns. Here, our results show that a restaurant can benefit from partnering with third-party delivery platforms when (i) the subscription fee is low, (ii) customers can flexibly decide whether to order from platforms or from restaurants directly, (iii) customers require more efficient delivery, (iv) average delivery distance is long, or (v) demand variance is high.

97 MATHEMATICS AND COMPUTING↗

The northeast materials database for magnetic materials

The discovery of magnetic materials with high operating temperature ranges and optimized performance is essential for advanced applications. Current data-driven approaches are limited by the lack of accurate, comprehensive, and feature-rich databases. This study aims to address this challenge by using Large Language Models (LLMs) to create a comprehensive, experiment-based, magnetic materials database named the Northeast Materials Database (NEMAD), which consists of 67,573 magnetic materials entries (www.nemad.org). The database incorporates chemical composition, magnetic phase transition temperatures, structural details, and magnetic properties. Enabled by NEMAD, we trained machine learning models to classify materials and predict transition temperatures. Our classification model achieved an accuracy of 90% in categorizing materials as ferromagnetic (FM), antiferromagnetic (AFM), and non-magnetic (NM). The regression models predict Curie (Néel) temperature with a coefficient of determination (R 2 ) of 0.87 (0.83) and a mean absolute error (MAE) of 56K (38K). These models identified 25 (13) FM (AFM) candidates with a predicted Curie (Néel) temperature above 500K (100K) from the Materials Project. This work shows the feasibility of combining LLMs for automated data extraction and machine learning models to accelerate the discovery of magnetic materials.

Ferromagnetism↗

Hot Droughts and Forest Tree Dynamics in the Amazon - Statistical Models, Scripts, Data, and Outputs

This package contains data, outputs, equations, and R scripts for analyses for manuscript entitled "Hot droughts in the Amazon: A window to a future hypertropical climate" by J. Chambers et al., in particular it contains statistical models and analyses for the INPA BIONTE tree mortality study. The Models folder contains details for all statistical models in PDF files. The Scripts folder contains the R scripts for Bayesian Hierarchical Models (two text files) and SEMs (one text file) are separate and reasonably annotated. All data associated with these scripts are in the data folder. The Data folder contains two of the three CSV files used for the analyses and are called by the R scripts. Two of them are part of published datasets (`BIONTE_mortality-rates.csv` from Lima et al. 2024, DOI:10.15486/ngt/1898910 and `SPEI.csv` from Pastorello et al. 2023 DOI:10.15486/ngt/1958257) and also provided in this package for convenience (please see the corresponding datasets for usage and citation terms). The third dataset (`BIONTE_gapfilled_wd.csv`) contains sensitive information and can be obtained by contacting the manuscript lead author. The Outputs folder contains the two output files that provide extra information about the analyses. The file `figuresFeb2025d.pdf` contains all the figures from the manuscript - captions are in the manuscript. The file `ChambersMS.pdf` contains primary results from Bayesian statistical models, regression analyses, and validation steps applied to the tree mortality data from the INPA experiments. The document includes visual summaries, model diagnostics, and leave-one-out (LOO) validation results. A breakdown of file contents can be found in the README file that is part of this package.

54 ENVIRONMENTAL SCIENCES↗

An Artificial Intelligence-Assisted Method for Dementia Detection Using Images from the Clock Drawing Test

Background: Widespread dementia detection could increase clinical trial candidates and enable appropriate interventions. Since the Clock Drawing Test (CDT) can be potentially used for diagnosing dementia-related disorders, it can be leveraged to develop a computer-aided screening tool. Objective: To evaluate if a machine learning model that uses images from the CDT can predict mild cognitive impairment or dementia. Methods: Images of an analog clock drawn by 3,263 cognitively intact and 160 impaired subjects were collected during in-person dementia evaluations by the Framingham Heart Study. We processed the CDT images, participant’s age, and education level using a deep learning algorithm to predict dementia status. Results: When only the CDT images were used, the deep learning model predicted dementia status with an area under the receiver operating characteristic curve (AUC) of 81.3% ± 4.3%. A composite logistic regression model using age, level of education, and the predictions from the CDT-only model, yielded an average AUC and average F1 score of 91.9% ±1.1% and 94.6% ±0.4%, respectively. Conclusion: Our modeling framework establishes a proof-of-principle that deep learning can be applied on images derived from the CDT to predict dementia status. When fully validated, this approach can offer a cost-effective and easily deployable mechanism for detecting cognitive impairment.

Neurosciences & Neurology↗

Relationships between Vehicle Pricing and Features: Data Driven Analysis of the Chinese Vehicle Market

A full-scale understanding of the dynamics of the Chinese vehicle market can benefit stakeholders with respect to rational decision-making and effective long-term investment. This study attempts to discover the common vehicle pricing patterns in the Chinese market by quantifying statistical correlations among critical vehicle features from intrinsic powertrain systems to extrinsic market positioning. The data samples involve almost all passenger vehicle models sold in 2013 to 2019. After comparing multiple statistical methodologies, a log-transformation variant of the multinomial linear regression model was found to be the best one, and the goodness of fit shows that this model can offer stable estimates, which were validated using 2019 market data. The insights achieved are: (1) The price and major performance features of SUVs/crossovers are similar to those of sedans; (2) If all other explicit features remain the same, the price of a Japanese midsize sedan is 62% higher than that of a Chinese midsize sedan, and European midsize vehicles have the highest prices overall. (3) The incremental price of fuel consumption varies by vehicle class and fuel economy. For example, from 30 to 50 MPG, the vehicle price increases by $119 for a Chinese brand sedan vehicle, by $69 for a Chinese brand SUV.

33 ADVANCED PROPULSION SYSTEMS↗

Application of machine learning approaches in the analysis of mass absorption cross-section of black carbon aerosols: Aerosol composition dependencies and sensitivity analyses

Physics-based models typically require an in-depth understanding of a phenomenon and assumptions of the underlying process(es), which are often hard to obtain in practice, whereas data-driven machine learning models learn the structure and patterns in the training data without any prior theoretical assumptions and then use inference to develop useful predictions. A novel machine learning-based algorithm has been previously developed for the prediction of black carbon mass absorption cross-section (MAC BC ) and applied to a variety of different atmospheric environments. In contrast to light scattering theories which require assumptions about the underlying physics, this algorithm uses time-series data of aerosol properties to estimate the temporally-varying MAC BC at 870 nm. Here, we analyze our algorithm and discuss the influence of aerosol optical properties (such as Ångström exponents and single scattering albedo) and chemical composition on the model outputs and the associated accuracy. Additionally, we conduct sensitivity analyses on our models to understand how the predictions change in response to different sets of input variables. Our support vector machine (SVM) for regression model is the least sensitive to variations in the input variables, although all models tend to exhibit a degradation to their accuracy when scattering Ångström exponents are less than one.

54 ENVIRONMENTAL SCIENCES↗

Artificial intelligence-based predictive modeling for imaging neutral particle analyzers on the DIII-D tokamak

The Imaging Neutral Particle Analyzer (INPA) at DIII-D is a diagnostic system used to accurately resolve the energy and spatial distributions of fast ions in fusion plasmas. A novel artificial intelligence (AI) technique named INPA-net is based on Reservoir Computing Networks and developed here to predict active and passive signals produced by charge-exchange reactions from injected and edge-cold neutrals, respectively, in magnetically confined fusion plasmas. This model is trained using a set of 21 time domain signals between 0 s to 3.35 s that includes injected beam and thermal plasma information, and 6444 real 2D experimental images of the INPA in 12 plasma discharges at DIII-D. The trained neural network is able to forecast experimental images in real-time. The model achieves an R-squared value of 0.91, which is higher than the 0.83 value achieved by a simple linear regression model. This improvement highlights the model's enhanced predictive accuracy for measured images from the validation set. This AI approach is valuable due to its rapid response times and potential for integration into real-time plasma control systems. A version of this model capable of generating syntehic images would be useful for the real-time monitoring of fast-ion transport. A comprehensive sensitivity study reveals that INPA-net maintains high performance even with variations in the input parameters, indicating the model's robustness and reliability. While developed for the INPA, the underlying architecture is adaptable and may be applied to various 2D imaging diagnostics in fusion research.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗