Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Model assignment”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

An Approach to Dynamic Human Reliability Analysis and Its Data Collection Framework

Human reliability analysis (HRA) is a method for evaluating human errors in a variety of complex systems such as nuclear power plants, military systems, aircraft, and chemical plants. Most HRA methods currently used by regulatory institutes or utilities are called static HRA and are carried out by simple worksheets or simple calculators. To date, there are many unsolved or intrinsic challenges in static HRA. For example, existing static HRA does not realistically model and evaluate human actions as they would be performed at actual systems. There is no method with HRA to objectively estimate the time required for human actions despite being essential to HRA processes. In addition, many HRA methods still rely on a dataset generated prior to the 1980s, from unrelated industry experience or simply from expert judgment. Accordingly, this study attempted to research how to overcome the challenges of existing HRA via dynamic risk assessment (a.k.a., simulation-based or computation-based risk assessment) techniques. First, this study developed a dynamic HRA method, named as PRocedure-based Investigation Method of EMRALD Risk Assessment – HRA (PRIMERA-HRA). The PRIMERA-HRA mainly concentrates on providing HRA analysts with specific guidelines on how to reasonably model human actions, assign human reliability data and evaluate output of simulation within a dynamic probabilistic risk assessment tool, called as Event Modeling Risk Assessment using Linked Diagrams (EMRALD). Second, this study also developed a module for performance shaping factors (i.e., the key concept in HRA quantification) applicable to dynamic HRA, then implemented it based on PRIMERA-HRA within the EMRALD tool. Third, this study developed an HRA data collection framework to support dynamic HRA, called as Simplified Human Error Experimental Program (SHEEP). Originally, the SHEEP study aimed to support static HRA and its data collection, but recently extended the scope to the new technologies such as dynamic HRA or HRA for advanced reactors. SHEEP focuses on the use of data collected from simplified simulators to complement—but not replace—data collection studies using full-scope simulators and actual operators. To date, many experiments were conducted under the SHEEP framework. Multiple analyses, such as human performance analysis, human error analysis, task complexity analysis, learning effect analysis and time distribution analysis, were also carried out using the collected data. Then, based on the major insights, an approach to inferring full-scope data based on simplified simulator data was proposed. The PRIMERA-HRA and SHEEP research are expected to evaluate human actions more realistically than existing static HRA, provide an opportunity to collect more HRA data with reasonable cost and labor, then contribute to enhance the quality of HRA.

99 - GENERAL AND MISCELLANEOUS↗

An Approach to Dynamic Human Reliability Analysis and Its Data Collection Framework

Human reliability analysis (HRA) is a method for evaluating human errors in a variety of complex systems such as nuclear power plants, military systems, aircraft, and chemical plants. Most HRA methods currently used by regulatory institutes or utilities are called static HRA and are carried out by simple worksheets or simple calculators. To date, there are many unsolved or intrinsic challenges in static HRA. For example, existing static HRA does not realistically model and evaluate human actions as they would be performed at actual systems. There is no method with HRA to objectively estimate the time required for human actions despite being essential to HRA processes. In addition, many HRA methods still rely on a dataset generated prior to the 1980s, from unrelated industry experience or simply from expert judgment. Accordingly, this study attempted to research how to overcome the challenges of existing HRA via dynamic risk assessment (a.k.a., simulation-based or computation-based risk assessment) techniques. First, this study developed a dynamic HRA method, named as PRocedure-based Investigation Method of EMRALD Risk Assessment – HRA (PRIMERA-HRA). The PRIMERA-HRA mainly concentrates on providing HRA analysts with specific guidelines on how to reasonably model human actions, assign human reliability data and evaluate output of simulation within a dynamic probabilistic risk assessment tool, called as Event Modeling Risk Assessment using Linked Diagrams (EMRALD). Second, this study also developed a module for performance shaping factors (i.e., the key concept in HRA quantification) applicable to dynamic HRA, then implemented it based on PRIMERA-HRA within the EMRALD tool. Third, this study developed an HRA data collection framework to support dynamic HRA, called as Simplified Human Error Experimental Program (SHEEP). Originally, the SHEEP study aimed to support static HRA and its data collection, but recently extended the scope to the new technologies such as dynamic HRA or HRA for advanced reactors. SHEEP focuses on the use of data collected from simplified simulators to complement—but not replace—data collection studies using full-scope simulators and actual operators. To date, many experiments were conducted under the SHEEP framework. Multiple analyses, such as human performance analysis, human error analysis, task complexity analysis, learning effect analysis and time distribution analysis, were also carried out using the collected data. Then, based on the major insights, an approach to inferring full-scope data based on simplified simulator data was proposed. The PRIMERA-HRA and SHEEP research are expected to evaluate human actions more realistically than existing static HRA, provide an opportunity to collect more HRA data with reasonable cost and labor, then contribute to enhance the quality of HRA.

99 - GENERAL AND MISCELLANEOUS↗

Time-Resolved X-ray Emission Spectroscopy and Synthetic High-Spin Model Complexes Resolve Ambiguities in Excited-State Assignments of Transition-Metal Chromophores: A Case Study of Fe-Amido Complexes

To fully harness the potential of abundant metal coordination complex photosensitizers, a detailed understanding of the molecular properties that dictate and control the electronic excited-state population dynamics initiated by light absorption is critical. In the absence of detectable luminescence, optical transient absorption (TA) spectroscopy is the most widely employed method for interpreting electron redistribution in such excited states, particularly for those with a charge-transfer character. The assignment of excited-state TA spectral features often relies on spectroelectrochemical measurements, where the transient absorption spectrum generated by a metal-to-ligand charge-transfer (MLCT) electronic excited state, for instance, can be approximated using steady-state spectra generated by electrochemical ligand reduction and metal oxidation and accounting for the loss of absorptions by the electronic ground state. However, the reliability of this approach can be clouded when multiple electronic configurations have similar optical signatures. Using a case study of Fe(II) complexes supported by benzannulated diarylamido ligands, we highlight an example of such an ambiguity and show how time-resolved X-ray emission spectroscopy (XES) measurements can reliably assign excited states from the perspective of the metal, particularly in conjunction with accurate synthetic models of ligand-field electronic excited states, leading to a reinterpretation of the long-lived excited state as a ligand-field metal-centered quintet state. Furthermore, a detailed analysis of the XES data on the long-lived excited state is presented, along with a discussion of the ultrafast dynamics following the photoexcitation of low-spin Fe(II)-N amido complexes using a high-spin ground-state analogue as a spectral model for the 5 T 2 excited state.

14 SOLAR ENERGY↗

Charging-management And Infrastructure-planning (cmip) Model

CMIP model explores various charging infrastructure network designs to serve a free-floating car-sharing fleet and determine the charging downtime experienced by the fleet for each design. Development of the CMIP model had two major steps: (1) describing modeling assumptions and (2) developing an integer program (IP) that jointly optimizes decisions about locations to install DC fast chargers and EV-to-charger assignments. The CMIP model integrates an EV charging model, EV energy consumption model, and heterogeneous, real-world vehicle use data with an integer programming optimization model to identify optimal location of new charging stations and calculate vehicle downtime for charging. The CMIP model can be applied to understand: (a) the reduction of EV fleet downtime if an additional fast-charging station is added to the current infrastructure and (b) to what extent total vehicle downtime would be sensitive to additional charging infrastructure.

Roni, MohammadS↗

Forecasting for ESCAPE: A Multi-Institution Hybrid Forecasting and Nowcasting Operation for Sea-Breeze Convection Supporting a Ground-Based and Airborne Field Campaign

The Experiment of Sea-Breeze Convection, Aerosols, Precipitation and Environment (ESCAPE) field project deployed two aircraft and ground-based assets in the vicinity of Houston, Texas, between 27 May and 2 July 2022, examining how meteorological conditions, dynamics, and aerosols control the initiation, early growth stage, and evolution of coastal convective clouds. To ensure that airborne- and ground-based assets were deployed appropriately, a forecasting and nowcasting team was formed. Daily forecasts guided real-time decision-making by assessing synoptic weather conditions, environmental aerosol, and a variety of atmospheric modeling data to assign a probability for meeting specific ESCAPE campaign objectives. During the research flights, a small team of forecasters provided “nowcasting” support by analyzing radar, satellite, and new model data in real time. The nowcasting team proved invaluable to the campaign operation, as sometimes changing environmental conditions affected, for example, the timing of convective initiation. In addition to the success of the forecasting and nowcasting teams, the ESCAPE campaign offered a unique “testbed” opportunity where in-person and virtual support both contributed to campaign objectives. The forecasting and nowcasting teams were each composed of new and experienced forecasters alike, where new forecasters were given invaluable experience that would otherwise be difficult to attain. Both teams received training on forecast models, map analysis, Hybrid Single-Particle Lagrangian Integrated Trajectory model (HYSPLIT), and thermodynamic sounding analysis before the beginning of the campaign. In this article, the ESCAPE forecasting and nowcasting teams reflect on these experiences, providing potentially useful advice for future field campaigns requiring forecasting and nowcasting support in a hybrid virtual/in-person framework.

54 ENVIRONMENTAL SCIENCES↗

EI_MS_ML

The unambiguous identification of compounds from their electron ionization mass (EI-MS) spectra remains a significant unsolved problem in the field of metabolomics and analytical chemistry as a whole. Typically EI-MS spectra are compared using various mathematical operations that convert the spectral similarity or differences into a distance-like metric that roughly approximates the similarity of any two spectra. A commonly used metric for this is the cosine similarity metric which has values close to one for very similar spectra and a value of zero for very dissimilar spectra; however, no metric is perfect. Due to the prevalence of structurally-similar compounds such as isomers and the prevalence of certain fragmentation patterns across structurally-dissimilar compounds, the unambiguous assignment of EI-MS spectra compounds remains difficult. Frequently, querying an observed EI-MS spectrum against a large database such as the NIST17 library yields multiple possible assignments requiring the end user to distinguish between multiple high scoring hits, or multiple low scoring hits while keeping in mind that the correct hit may not be in the database at all. Although techniques such as orthogonal information from techniques such as chromatography can greatly aid in unambiguous assignment, this also requires more complicated experimental designs and access to more complicated analytical instrumentation. Substructures can be trivially detected and represented as strings using a previously published technique called node coloring from a known chemical structure. However, for experimentally-derived EI-MS spectra this information must be derived from the spectra itself (i.e., because we do not know what compound it represents). To achieve this, the software uses techniques from the field of machine learning and a large training dataset of EI-MS spectra corresponding to known structures annotated with substructure strings, to build models that can predict the presence of a given chemical substructure from an EI-MS spectrum directly.If these predictions are of high-quality (i.e., are unlikely to be false positives), the presence of one or more predicted substructures can be used to constrain the number of possible hits for a query spectrum. Mathematically, this restriction could be expressed in many forms, but the most straight-forward implementation is to weight the cosine similarity of a query spectrum and a plausible database match with a Tanimoto-like coefficient based on the ratio of the number of substructures predicted to the number of substructures present in the potential database hit. Determining which combination of models best reduces assignment ambiguity will be achieved using a combination of manual curation and optimization techniques such as genetic algorithms. This software will perform all the steps necessary to construct said models from a training dataset and evaluate them using a holdout dataset. Various statistical analyses can be performed to determine if this approach does decrease assignment ambiguity. For example, if this approach works, on average, the rank-order of the correct assignment for the holdout set of EI-MS spectra should decrease and the weighted cosine similarities for most of the possible matches in the database should be better than the unweighted cosine similarities. Furthermore, this same pipeline can be used on real experimental data to generate less ambiguous assignments.

Mitchell, Joshua↗

Structure Prediction of Ionic Epitaxial Interfaces with Ogre Demonstrated for Colloidal Heterostructures of Lead Halide Perovskites

Colloidal epitaxial heterostructures are nanoparticles composed of two different materials connected at an interface, which can exhibit properties different from those of their individual components. Combining dissimilar materials offers exciting opportunities to create a wide variety of functional heterostructures. However, assessing structural compatibility–the main prerequisite for epitaxial growth–is challenging when pairing complex materials with different lattice parameters and crystal structures. This complicates both the selection of target heterostructures for synthesis and the assignment of interface models when new heterostructures are obtained. Here, we demonstrate Ogre as a powerful tool to accelerate the design and characterization of colloidal heterostructures. To this end, we implemented developments tailored for the high-efficiency prediction of epitaxial interfaces between ionic/polar materials, which encompass most colloidal semiconductors. These include the use of pre-screening candidate models based on charge balance at the interface and the use of a classical potential for fast energy evaluations, with parameters automatically calculated based on the input bulk structures. These developments are validated for perovskite-based CsPbBr 3 /Pb 4 S 3 Br 2 heterostructures, where Ogre produces interface models in excellent agreement with density functional theory and experiments. Furthermore, we use Ogre to rationalize the templating effect of CsPbCl 3 on the growth of lead sulfochlorides, where perovskite seeds induce the formation of Pb 4 S 3 Cl 2 rather than Pb 3 S 2 Cl 2 due to better epitaxial compatibility. Finally, combining Ogre simulations with experimental data enables us to unravel the structure and composition of the hitherto unsolved CsPbBr 3 /Bi x Pb y S z interface, and to assign a structure to several other reported metal halide- and oxide-based interfaces. The Ogre package is available on GitHub or via the OgreInterface desktop application, available for Windows, Linux, and Mac.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Model Package Report: Central Plateau Vadose Zone Models

This model package report describes concisely the modeling objectives, conceptualization, implementation, uncertainty and sensitivity, configuration control, limitations, data needs, and recommendations for improvement of the Central Plateau Vadose Zone Models. This collection of models is developed to meet the vadose zone simulation needs of the Hanford Site Composite Analysis (CA), the Hanford Site Cumulative Impact Evaluation (CIE), and are expected to find other applications in Hanford Site remedial cleanup decision-making processes. This model package report describes and documents the development of the models themselves to fulfill technical approaches defined for the CA and the CIE. This report does not document any specific calculation using these models: instead, applications of individual vadose models to perform specific calculations will be documented in environmental calculation files, including inputs and results, as appropriate. Configuration management for the Central Plateau Vadose Zone Models is managed through assignment of unique sequential numerical version numbers and archival of the models in the Environmental Modeling Management Archive.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

Identification of a new isomeric state in 76 Zn following the β decay of 76 Cu

Background: The evolution of nuclear shell structure far from stability can be explored by identifying and measuring the properties of isomers. Neutron-rich nuclei between the Z = 28 and Z = 50 closed shells have been the subject of recent studies which have identified a number of 0.1 - 10 µs isomers and measured detailed spectroscopic properties. Purpose: The purpose of this analysis was to identify and measure the properties of short-lived isomeric states populated following β decay in Z ≈ 30, N ≈ 50 nuclei near the doubly magic nucleus 78 Ni. Methods: Here, radioactive ions produced by beam fragmentation at the National Superconducting Cyclotron Laboratory were implanted into a CeBr 3 scintillator coupled to a pixelated photomultiplier tube. Ancillary arrays of HPGe clover and LaBr 3 detectors were positioned around the implantation detector to measure β-delayed γ rays. Results: The previously observed 2634-keV level in 76 Zn, populated following the β decay of 76 Cu, was identified as isomeric with a half-life of 25.4(4) ns. A combination of timing and γ-ray spectroscopy was used to confirm this assignment. Shell-model calculations were performed and indicate that this state may be a high-spin negativeparity state formed by the occupation of the ν0g 9/2 orbital. Conclusions: A new isomeric state in 76 Zn has been identified and its half-life was measured. Ambiguity about the structure of this state could be resolved with further experiments.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Breakup dynamics in a pressure-swirl injector for urea-water solution applications: A computational study

The co-optimization of in-cylinder combustion and after-treatment technology has become a major aspect in engine design and development, with the goal of meeting the increasingly restrictive emission regulations in the transportation industry. Selective Catalytic Reduction is a robust technology to control the emission of NO x , and the injection of urea in water solution is the exhaust tailpipe is a key aspect of its operation. The proposed work uses high-fidelity Computational Fluid Dynamics to characterize the atomization dynamics of the liquid jet in relevant cross-flow conditions. The study focuses on a commercial low-pressure (9 bar) pressure-swirl injector which is characterized in its internal geometry through high-resolution X-ray micro-computational tomography. The internal two-phase flow has been modeled according to the volume-of-fluid approach in a large eddy simulation framework and validated against near-nozzle X-ray radiography measurement. Moreover, characterizing the breakup dynamics for the swirling hollow cone formation, and assessing the influence of the cross-flow in the breakup dynamics was completed. The results have been reported proposing Re-Oh maps and probability density functions of the spray kinematics. Higher cross-flow momentum generates an increase in the jet intact length and a reduction of the liquid droplet diameters. The axial momentum of the jet is affected by the cross-flow already in the near-nozzle region, determining a relevant deviation of the spray velocities. In conclusion, this work aims to inform the initialization of Eulerian-Lagrangian spray models through the assignment of droplet kinematics and static one-way coupling between volume-of-fluid results and Lagrangian spray parcels, to be used for system-size domain simulations.

33 ADVANCED PROPULSION SYSTEMS↗

Attract-repel path planner system for collision avoidance

A system for determining a travel direction that avoids objects when a vehicle travels from a current location to a target location is provided. The system determines a travel direction based on an attract-repel model. The system assigns a repel value to the object locations and an attract value. A repel represents a magnitude of a directional repulsive force, and the attract value represents the magnitude of a directional repulsive force. The system calculates an attract-repel field having an attract-repel magnitude and attract-repel direction for the current location based on the repel values and their directions and the attract value and its direction. The system then determines the travel direction for a vehicle to be the direction of the attract-repel field at the current location.

Paglieroni, David W.↗

Hydrostratigraphic Region 1 Model

Farnsworth Unit (FWU) CO2EOR Eclipse compositional model: Hydrostratigraphic Region 1 model uses the HS1 relative permeability curves assigned heterogeneously by hydrostratigraphic unit. The model with capillary pressure applies the HS1 curves heterogeneously by hydrostratigraphic unit.

Capillary Pressure↗

Hydrostratigraphic Region 3 Model - with and without HSU Capillary Pressure

Farnsworth Unit (FWU) CO2EOR Eclipse compositional model: Hydrostratigraphic Region 3 model uses the HS3 relative permeability curves assigned heterogeneously by hydrostratigraphic unit. The model with capillary pressure applies the HS3 curves heterogeneously by hydrostratigraphic unit.

Capillary Pressure↗

Assessing mechanical response of CO 2 storage into a depleted carbonate reef using a site-scale geomechanical model calibrated with field tests and InSAR monitoring data

Geomechanical risks of injection have raised concerns regarding secure CO 2 storage. In this work, a combined monitoring and modeling approach is used to assess the stress changes and surface uplift associated with CO 2 injection into a depleted carbonate reef of the Michigan basin. A site-scale geomechanical model is built by assigning mechanical properties of formations using well-log and experimental data. Gravity load is applied to the model to estimate the vertical component of stress as well as different lateral boundary displacement scenarios to estimate horizontal stresses. We used a poroelastic pressure-dependent model (instead of a linear elastic mechanical earth model) to calibrate initial stresses using hydraulic fracture test data measured at depleted reservoir status. Multi-phase fluid flow-geomechanical simulations are performed to estimate the poroelastic response during (1) primary depletion (2) field-scale CO 2 injection phase (3) a hypothetical forecast scenario in which well bottom hole pressure (BHP) reach 45000 KPa. The predicted surface uplift is less than 1 mm at the end of the field-scale CO 2 injection phase which is in good agreement with Interferometric Synthetic Aperture Radar (InSAR) uplift measurement. Although the InSAR data shows an insignificant uplift, hydromechanical modeling of injection shows that CO 2 injection still causes reservoir deformation emphasizing the role of carbonate overburden and reservoir formation mechanical properties and limited size of reef on diminishing the surface deformation. Modeling indicates poroelastic response of caprock matters to estimate uplift. The lower permeability of the top two layers provides additional barrier to large uplift. Also, history of subsidence due to production should be accounted to predict uplift due to a follow up injection correctly. This report shows the significance of combining a calibrated geomechanical model with field measured stresses and monitoring data to be used as a tool to ensure the safety of CO 2 storage.

42 ENGINEERING↗

Determining Partial Atomic Charges for Liquid Water: Assessing Electronic Structure and Charge Models

Partial atomic charges provide an intuitive and efficient way to describe the charge distribution and the resulting intermolecular electrostatic interactions in liquid water. Many charge models exist and it is unclear which model provides the best assignment of partial atomic charges in response to the local molecular environment. In this work, we systematically scrutinize various electronic structure methods and charge models (Mulliken, natural population analysis, CHelpG, RESP, Hirshfeld, Iterative Hirshfeld, and Bader) by evaluating their performance in predicting the dipole moments of isolated water, water clusters, and liquid water as well as charge transfer in the water dimer and liquid water. Although none of the seven charge models is capable of fully capturing the dipole moment increase from isolated water (1.85 D) to liquid water (about 2.9 D), the Iterative Hirshfeld method performs best for liquid water, reproducing its experimental average molecular dipole moment, yielding a reasonable amount of intermolecular charge transfer, and showing modest sensitivity to the local water environment. The performance of the charge model is dependent on the choice of the density functional and the quantum treatment of the environment. The computed molecular dipole moment of water generally increases with the percentage of the exact Hartree–Fock exchange in the functional, whereas the amount of charge transfer between molecules decreases. For liquid water, including two full solvation shells of surrounding water molecules (within about 5.5 Å of the central water) in the quantum chemical calculation converges the charges of the central water molecule. Furthermore, our final pragmatic quantum chemical charge-assigning protocol for liquid water is the Iterative Hirshfeld method with M06-HF/aug-cc-pVDZ and a quantum region cutoff radius of 5.5 Å.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A Compound Poisson Generator Approach to Point-source Inference in Astrophysics

Abstract The identification and description of point sources is one of the oldest problems in astronomy, yet even today the correct statistical treatment for point sources remains one of the field’s hardest problems. For dim or crowded sources, likelihood-based inference methods are required to estimate the uncertainty on the characteristics of the source population. In this work, a new parametric likelihood is constructed for this problem using compound Poisson generator (CPG) functionals that incorporate instrumental effects from first principles. We demonstrate that the CPG approach exhibits a number of advantages over non-Poissonian template fitting (NPTF)—an existing method—in a series of test scenarios in the context of X-ray astronomy. These demonstrations show that the effect of the point-spread function, effective area, and choice of point-source spatial distribution cannot, generally, be factorized as they are in NPTF, while the new CPG construction is validated in these scenarios. Separately, an examination of the diffuse-flux emission limit is used to show that most simple choices of priors on the standard parameterization of the population model can result in unexpected biases: when a model comprising both a point-source population and diffuse component is applied to this limit, nearly all observed flux will be assigned to either the population or to the diffuse component. A new parameterization is presented for these priors that properly estimates the uncertainties in this limit. In this choice of priors, CPG correctly identifies that the fraction of flux assigned to the population model cannot be constrained by the data.

79 ASTRONOMY AND ASTROPHYSICS↗

Water Network Tool for Resilience (WNTR)

The Water Network Tool for Resilience (WNTR) is an open source Python package designed to simulate and analyze resilience of water distribution networks. The United States Environmental Protection Agency, in partnership with Sandia National Laboratories, developed WNTR to integrate critical aspects of resilience modeling for water distribution networks into a single software framework. The software includes capability to: • Generate water network models • Modify network structure and operations • Assign fragility and survival curves to network components • Model disruptive events such as power outages, earthquakes, fires, pipe breaks, and contamination incidents • Model response and repair strategies • Simulate hydraulics and water quality • Evaluate resilience using a wide range of metrics • Integrate dependency with other critical infrastructure and supply chains • Analyze results and generate graphics SAND2019-450 M

Villa, Daniel↗

Ch3MS-RF: a random forest model for chemical characterization and improved quantification of unidentified atmospheric organics detected by chromatography–mass spectrometry techniques

Abstract. The chemical composition of ambient organic aerosols plays a critical role in driving their climate and health-relevant properties and holds important clues to the sources and formation mechanisms of secondary aerosol material. In most ambient atmospheric environments, this composition remains incompletely characterized, with the number of identifiable species consistently outnumbered by those that have no mass spectral matches in the literature or the National Institute of Standards and Technology/National Institutes of Health/Environmental Protection Agency (NIST/NIH/EPA) mass spectral databases, making them nearly impossible to definitively identify. This creates significant challenges in utilizing the full analytical capabilities of techniques which separate and generate spectra for complex environmental samples. In this work, we develop the use of machine learning techniques to quantify and characterize novel, or unidentifiable, organic material. This work introduces Ch3MS-RF (Chemical Characterization by Chromatography–Mass Spectrometry Random Forest Modeling), an open-source, R-based software tool, for efficient machine-learning-enabled characterization of compounds separated in chromatography–mass spectrometry applications but not identifiable by comparison to mass spectral databases. A random forest model is trained and tested on a known 130 component representative external standard to predict the response factors of novel environmental organics based on position in volatility–polarity space and mass spectrum, enabling the reproducible, efficient, and optimized quantification of novel environmental species. Quantification accuracy on a reserved 20 % test set randomly split from the external standard compound list indicates that random forest modeling significantly outperforms the commonly used methods in both precision and accuracy, with a median response factor percent error of −2 %, for modeled response factors, compared to > 15 %, for typically used proxy assignment-based methods. Chemical properties modeling, evaluated on the same reserved 20 % test set and an extrapolation set of species identified in ambient organic aerosol samples collected in the Amazon rainforest, also demonstrate robust performance. Extrapolation set property prediction mean absolute errors for carbon number, oxygen to carbon ratio (O : C), average carbon oxidation state (OSc‾), and vapor pressure are 1.8, 0.15, 0.25, and 1.0 (log(atm)), respectively. Extrapolation set out-of-sample R2 for all properties modeled are above 0.75, with the exception of vapor pressure. While predictive performance for vapor pressure is less robust compared to the other chemical properties modeled, random-forest-based modeling was significantly more accurate than other commonly used methods of vapor pressure prediction, decreasing the mean vapor pressure prediction error to 0.24 (log(atm)) from 0.55 (log(atm)) (chromatography-based vapor pressure prediction) and 1.2 (log(atm)) (chemical formula-based vapor pressure prediction). The random forest model significantly advances an untargeted analysis of the full scope of chemical speciation yielded by two-dimensional gas chromatography (GCxGC-MS) techniques and can be applied to gas chromatography coupled with electron ionization mass spectrometry (GC-MS) as well. It enables the accurate estimation of key chemical properties commonly utilized in the atmospheric chemistry community, which may be used to more efficiently identify important tracers for further individual analysis and to characterize compound populations uniquely formed under specific ambient conditions.

54 ENVIRONMENTAL SCIENCES↗