Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “large data sets”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Molecular Modeling and Molecular Dynamics Simulation of a Packed and Intact Bacterial Microcompartment

Bacterial microcompartments (BMCs) are protein-bound organelles found in some bacteria which encapsulate enzymes for enhanced catalytic activity. These compartments spatially sequester enzymes within semipermeable shell proteins and are packed full of enzyme cargoes and metabolites as they fulfill their function. Coupling together recent SAXS and proteomics work, it is possible to develop molecular models for these microcompartments and interrogate enzyme and metabolite dynamics within. Our primary goal of this study is to quantify the permeability of metabolite glyceraldehyde-3-phosphate (G3P) and dihydroxyacetone phosphate (DHAP) across the BMC shell through classical molecular dynamics simulation. The Haliangium ochraceum model of BMC shell (PDB: 6MZX) was used to model an intact BMC of approximately 10 million atoms. Working at this scale presented its own challenges in managing large data sets, with multiple challenges and hardware advances discussed that facilitated this work. Over approximately 750 ns of aggregate simulation, we see multiple permeation events for these metabolites that were added at high concentration through the pores present within BMC shell tiles. When compared to independent permeability estimates for the same metabolites determined through replica exchange umbrella sampling simulations, the permeabilities varied by approximately 3 orders of magnitude. Regardless, the permeability coefficients for both G3P and DHAP are highly similar and very high, such that only very small concentration gradients can be maintained across the BMC shell between the cytosol and BMC interior. The large simulation systems also facilitated comparisons for molecular diffusivity in the crowded environment within the BMC shell. By our estimates, the viscosity within a packed BMC shell is at least 10-fold higher than it would be in neat solution and is the real driver for varying permeability estimates we obtained through simulation. These findings will be used as design inputs for future bioengineering efforts to make products from BMCs, highlighting how permeable BMC shells can be.

Diffusion↗

Generative AI for Grid Operations [Slides]

In the last few years, the development and use of generative artificial intelligence (AI) and large-language models (LLMs) have changed the landscape of how AI and machine learning (ML) are being used in power systems. LLMs are built on foundational models based on large data sets that can be trained to provide information rapidly and through simple natural language prompts. Generative AI can then perform human-like tasks using ML models to identify and mimic pattens in the data sets. This presentation explores how generative AI can enhance grid operations by improving forecasts, enabling rapid contingency analyses, and offering real-time operational suggestions. By providing grid operators with valuable insights, generative AI will empower them to manage power systems more effectively.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Towards Automating Spacecraft Attitude Sensor Calibration

With a view towards reducing cost and complexity for spacecraft early mission support at the NASA Goddard Space Flight Center (GSFC), efforts are being made to automate the attitude sensor calibration process. This paper addresses one of the major components needed by such a system. The beneficiaries of an improved calibration process are missions that demand moderate to high precision attitude knowledge or that need to perform accurate attitude slews. Improved slew accuracy reduces the time needed for re-acquisition of fine-pointing after each attitude maneuver, Rapid target acquisition can be very important for astronomical targeting or for off-nadir surface feature targeting by Earth-oriented spacecraft. The normal sequence of on-orbit calibration starts with alignment calibration of the star trackers and possibly the Sun sensor. Their relative alignment needs to be determined using a sufficiently large data set so their fields of view are adequately sampled. Next, the inertial reference unit (IRU) is calibrated for corrections to its alignment and scale factors. The IRU biases are estimated continuously by the onboard attitude control system, but the IRU alignment and scale factors are usually determined on the ground using a batch-processing method on a data set that includes several slews sufficient to give full observability of all the IRU calibration parameters. Finally, magnetometer biases, alignment, and its coupling to the magnetic torquers are determined in order io improve momentum management and occasionally for use in the attitude determination system. The detailed approach used for automating calibrations will depend on whether the automated system resides on the ground or on the spacecraft with an ultimate goal of autonomous calibration. Current efforts focus on a ground-based system driving subsystems that could run either on the ground or onboard. The distinction is that onboard calibration should process the data sequentially rather than in a single large batch since onboard computer data storage is limited. Very good batch- processing calibration utilities have been developed and used extensively at NASA/GSFC for mission support but no sequential calibration utilities are available. To meet this need, this paper presents the mathematical description of a sequential IRU calibration system. The system has been tested using flight data from the Rossi X-ray Timing Explorer (RXTE) during a series of attitude slews. The paper also discusses the current state of the overall automated system and describes plans for adding sequential alignment calibration and other additions that will reduce the amount of analyst time and input.

Sedlak, Joseph↗

Scalable probabilistic estimates of electric vehicle charging given observed driver behavior

To prepare for rapid growth in global electric vehicle adoption, grid and policy planners depend on detailed forecasts of future charging demand. In this paper we propose a novel holistic, scalable, probabilistic framework to produce large-scale estimates of electric vehicle charging load for long-term planning that capture real drivers’ charging patterns. Our framework captures the uncertainty and stochasticity in charging demand by taking a graphical modeling approach. It has three core elements: driver groups, charging segment choices, and charging session time and energy requirements. The framework uses hierarchical clustering to group drivers by their charging histories, capturing their heterogeneous behaviors and preferences across different segments or types of charging. The framework uses probabilistic mixture models for each driver group’s sessions to identify the unique charging behaviors observed within each segment. We illustrate its application with a large data set from California, profiling the charging patterns and unique driver clusters it identifies. Using the model knobs representing drivers’ battery capacities, behavior, and segment access we present scenarios for California’s charging demand in 2030 with 8 million passenger electric vehicles. Peak charging demand ranged from 3.3 to 8.7 GW across scenarios. Furthermore, each was calculated in under 45 s on a laptop computer.

33 ADVANCED PROPULSION SYSTEMS↗

Predicting the viability of beta-lactamase: How folding and binding free energies correlate with beta-lactamase fitness

One of the long-standing holy grails of molecular evolution has been the ability to predict an organism’s fitness directly from its genotype. With such predictive abilities in hand, researchers would be able to more accurately forecast how organisms will evolve and how proteins with novel functions could be engineered, leading to revolutionary advances in medicine and biotechnology. In this work, we assemble the largest reported set of experimental TEM-1 β-lactamase folding free energies and use this data in conjunction with previously acquired fitness data and computational free energy predictions to determine how much of the fitness of β-lactamase can be directly predicted by thermodynamic folding and binding free energies. We focus upon β-lactamase because of its long history as a model enzyme and its central role in antibiotic resistance. Based upon a set of 21 β-lactamase single and double mutants expressly designed to influence protein folding, we first demonstrate that modeling software designed to compute folding free energies such as FoldX and PyRosetta can meaningfully, although not perfectly, predict the experimental folding free energies of single mutants. Interestingly, while these techniques also yield sensible double mutant free energies, we show that they do so for the wrong physical reasons. We then go on to assess how well both experimental and computational folding free energies explain single mutant fitness. We find that folding free energies account for, at most, 24% of the variance in β-lactamase fitness values according to linear models and, somewhat surprisingly, complementing folding free energies with computationally-predicted binding free energies of residues near the active site only increases the folding-only figure by a few percent. This strongly suggests that the majority of β-lactamase’s fitness is controlled by factors other than free energies. Overall, our results shed a bright light on to what extent the community is justified in using thermodynamic measures to infer protein fitness as well as how applicable modern computational techniques for predicting free energies will be to the large data sets of multiply-mutated proteins forthcoming.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

SAXS Assistant: Automated SAXS analysis for structural discovery in biologics and polymeric nanoparticles

Small-angle x-ray scattering (SAXS) is a powerful technique for assessing macromolecular structure. High-throughput SAXS is limited by the time-consuming and, at times, subjective nature of SAXS data interpretation. Here, we present SAXS Assistant, a Python-based script that streamlines SAXS data analysis to extract features for machine learning (ML) and key structural parameters, including the Guinier radius of gyration (R g ), pair distance distribution function (PDDF)-derived R g , maximum particle dimension (D max ), and Kratky plots. The script builds upon BioXTAS RAW and validates reliability via Guinier/PDDF R g agreement, an important indicator of well-measured data sets. For assistance in D max estimation, a multilayer perceptron regressor was trained with 1940 data files from the Small Angle Scattering Biological Data Bank. The model achieved a test set performance R 2 = 0.90 and mean absolute error = 11.7 Å. Training exclusively with experimental data translates analyses from researchers, including experts in the field, to the ML model, which helps assess D max estimations from PDDF. Gaussian mixture model clustering was implemented to classify profiles into structural classes based on entries in the Small Angle Scattering Biological Data Bank. Users may therefore assess the similarity between experimental samples and known biomolecular shapes within the mapped repository entries. This probabilistic clustering aids in quantifying information from Kratky and generating shape-descriptive features. SAXS Assistant accelerates SAXS data analysis through enforced quality control, ML-ready outputs, and flags for low-confidence results. In addition to providing the ability to analyze large data sets at high throughput, this tool is versatile and may serve researchers in both biological and synthetic polymer research fields.

36 MATERIALS SCIENCE↗

Deep-Learning-Based Segmentation of Keyhole in In-Situ X-ray Imaging of Laser Powder Bed Fusion

In laser powder bed fusion processes, keyholes are the gaseous cavities formed where laser interacts with metal, and their morphologies play an important role in defect formation and the final product quality. The in-situ X-ray imaging technique can monitor the keyhole dynamics from the side and capture keyhole shapes in the X-ray image stream. Keyhole shapes in X-ray images are then often labeled by humans for analysis, which increasingly involves attempting to correlate keyhole shapes with defects using machine learning. However, such labeling is tedious, time-consuming, error-prone, and cannot be scaled to large data sets. To use keyhole shapes more readily as the input to machine learning methods, an automatic tool to identify keyhole regions is desirable. In this paper, a deep-learning-based computer vision tool that can automatically segment keyhole shapes out of X-ray images is presented. The pipeline contains a filtering method and an implementation of the BASNet deep learning model to semantically segment the keyhole morphologies out of X-ray images. The presented tool shows promising average accuracy of 91.24% for keyhole area, and 92.81% for boundary shape, for a range of test dataset conditions in Al6061 (and one AliSi10Mg) alloys, with 300 training images/labels and 100 testing images for each trial. Prospective users may apply the presently trained tool or a retrained version following the approach used here to automatically label keyhole shapes in large image sets.

36 MATERIALS SCIENCE↗

Comparison of Equilibrium and Nonequilibrium Approaches for Relative Binding Free Energy Predictions

Alchemical relative binding free energy calculations have recently found important applications in drug optimization. A series of congeneric compounds are generated from a preidentified lead compound, and their relative binding affinities to a protein are assessed in order to optimize candidate drugs. While methods based on equilibrium thermodynamics have been extensively studied, an approach based on nonequilibrium methods has recently been reported together with claims of its superiority. However, these claims pay insufficient attention to the basis and reliability of both methods. Here we report a comparative study of the two approaches across a large data set, comprising more than 500 ligand transformations spanning in excess of 300 ligands binding to a set of 14 diverse protein targets. Ensemble methods are essential to quantify the uncertainty in these calculations, not only for the reasons already established in the equilibrium approach but also to ensure that the nonequilibrium calculations reside within their domain of validity. If and only if ensemble methods are applied, we find that the nonequilibrium method can achieve accuracy and precision comparable to those of the equilibrium approach. Compared to the equilibrium method, the nonequilibrium approach can reduce computational costs but introduces higher computational complexity and longer wall clock times. There are, however, cases where the standard length of a nonequilibrium transition is not sufficient, necessitating a complete rerun of the entire set of transitions. This significantly increases the computational cost and proves to be highly inconvenient during large-scale applications. Our findings provide a key set of recommendations that should be adopted for the reliable implementation of nonequilibrium approaches to relative binding free energy calculations in ligand-protein systems.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Space Launch System Liftoff and Separation Dynamics Analysis Tool Chain

A flexible, hierarchical tool chain that is being applied to NASA’s Space Launch System (SLS) for critical dynamics phenomena is described. This tool chain, called CLVTOPS, is used to investigate lateral liftoff movement of the vehicle as it departs and clears the mobile launch tower and separation of the two solid rocket boosters without collision with the core stage and payload. The toolset’s architecture was configured to take advantage of a modern software-engineering approach for maximum flexibility and utilization of open-source simulations and associated tools. As opposed to a “monolithic” approach, scripting languages were used to “bind” together a tool chain to configure and organize input data, execute and produce analysis results, and post-process these results to facilitate a rapid iterative analysis process to quickly address issues and pursue alternatives with emphasis on analysis automation. Key capabilities in the tool chain include processing and mining of very large data sets, a wide range of graphical depictions, and high-fidelity, physics-based simulations. The paper begins with a problem description and the motivation for liftoff and separation dynamics analysis followed by a historical survey of dynamics analyses for previous NASA human-rated launch vehicles. Details of the tool chain and its components are then introduced divided, first, into description of the scripting language architecture used to “bind” the simulation tools, programs, and scripts together and, second, the physics models and simulations. Representative analyses and data products are shown for liftoff and booster separation dynamics that provide in-depth insight to the tool chain’s capabilities. Supporting activities such as simulation tool chain verification, version archiving and data management, and training are addressed. The paper concludes with case examples on how the tool chain can be tailored to related aerospace dynamics analyses, both large and small. These patterns and techniques for SLS dynamics tool construction can be applied for other aerospace simulations.

6DOF↗

Space Launch System Liftoff and Separation Dynamics Analysis Tool Chain

A flexible, hierarchical tool chain that is being applied to NASA’s Space Launch System (SLS) for critical dynamics phenomena is described. This tool chain, called CLVTOPS, is used to investigate lateral liftoff movement of the vehicle as it departs and clears the mobile launch tower and separation of the two solid rocket boosters without collision with the core stage and payload. The toolset’s architecture was configured to take advantage of a modern software engineering approach for maximum flexibility and utilization of open-source simulations and associated tools. As opposed to a “monolithic” approach, scripting languages were used to “bind” together a tool chain to configure and organize input data, execute and produce analysis results, and post-process these results to facilitate a rapid, iterative analysis process to quickly address issues and pursue alternatives with emphasis on analysis automation. Key capabilities in the tool chain include processing and mining of very large data sets, a wide range of graphical depictions, and high-fidelity, physics-based simulations. The paper begins with a problem description and the motivation for liftoff and separation dynamics analysis followed by a historical survey of dynamics analyses for previous NASA human-rated launch vehicles. Details of the tool chain and its components are then introduced and divided, first, into description of the scripting language architecture used to “bind” the simulation tools, programs, and scripts together and, second, the physics models and simulations. Representative analyses and data products for liftoff and booster separation dynamics are shown in order to provide in-depth insight into the tool chain’s capabilities. Supporting activities such as simulation tool chain verification, version archiving and data management, and training are addressed. The paper concludes with case examples on how the tool chain can be tailored to related aerospace dynamics analyses, both large and small. The flexibility and versatility of this tool chain in supporting analyses of such a diverse range of aerospace applications demonstrates the feasibility of applying these patterns and techniques for tool construction to other aerospace simulations.

6DOF↗

Improving the Automatic Inversion of Digital ISIS-2 Ionogram Reflection Traces into Topside Vertical Electron-Density Profiles

The topside-sounders on the four satellites of the International Satellites for Ionospheric Studies (ISIS) program were designed as analog systems. The resulting ionograms were displayed on 35-mm film for analysis by visual inspection. Each of these satellites, launched between 1962 and 1971, produced data for 10 to 20 years. A number of the original telemetry tapes from this large data set have been converted directly into digital records. Software, known as the TOPside Ionogram Scalar with True-height (TOPIST) algorithm has been produced that enables the automatic inversion of ISIS-2 ionogram reflection traces into topside vertical electron-density profiles Ne(h). More than million digital Alouette/ISIS topside ionograms have been produced and over 300,000 are from ISIS 2. Many of these ISIS-2 ionograms correspond to a passive mode of operation for the detection of natural radio emissions and thus do not contain ionospheric reflection traces. TOPIST, however, is not able to produce Ne(h) profiles from all of the ISIS-2 ionograms with reflection traces because some of them did not contain frequency information. This information was missing due to difficulties encountered during the analog-to-digital conversion process in the detection of the ionogram frame-sync pulse and/or the frequency markers. Of the many digital topside ionograms that TOPIST was able to process, over 200 were found where direct comparisons could be made with Ne(h) profiles that were produced by manual scaling in the early days of the ISIS program. While many of these comparisons indicated excellent agreement (<10% average difference over the entire profile) there were also many cases with large differences (more than a factor of two). Here we will report on two approaches to improve the automatic inversion process: (1) improve the quality of the digital ionogram database by remedying the missing frequency-information problem when possible, and (2) using the above-mentioned comparisons as teaching examples of how to improve the original TOPIST software.

Benson, R. F.↗

What Can We Learn from One Billion Ground System Log Messages?

Shortage of log-based data in a ground system they have traditionally been the under achievers in a satellite ground system. This is due to several factors: Once log messages scroll out of view on the TTC event console window they are soon forgotten. Application and system log files are scattered across directories within a system, across a multitude of servers, and across one or more databases making access cumbersome. Typical tools to perform log file content searching are generally crude and typically only employed as part of trouble-shooting exercises.As we move towards satellite constellations and fleets and add even more status information, the number of messages keeps growing. One mission now estimates that they could generate 3,000,000 messages per day 1 billion per year - for the life of their mission. What to do with those 1 billion messages? That is the challenge. With the recent technological advances in the management of large data sets, text-based processing, and data analytics, there are now capabilities that we can provide to the ground system engineers and satellite operators to address what we postulate are missed opportunities. Advanced real-time log analysis can allow us to be less reactionary in favor of being more proactive. Analytics goals include the ability to: Identify root cause of unexpected events, failures or error conditions enabled by correlating disparate data. Detect security breaches attempts before they are successful. Help admins ensure IT resources continue running optimally. Identify trends and patterns that may indicate impending failures or error conditions for valuable assets before they happen. Compare satellites in a fleet or constellation in terms of number of alarms, number of command sent to them, etc.. Answer questions like "Are the operations support needs increasing over the past year?" or "Have we seen this combination of alarm conditions before?" But really, once the tools are readily available the users will start realizing what can be done with their new powers. In this presentation we will show the results of analyzing millions of actual mission operations log messages, how the results can be displayed to the user, and how new products now available as open source can be applied to the challenges of large scale time-tagged text-based mission operations messages. Flight operations team members believe that this is a powerful new option for how they assess overall system and space asset health. Technical descriptions of the design, tools, and storage will be provided. One billion messages? Bring'em on!

Orsborne, Sharon↗

Mass Dependence of Galaxy–Halo Alignment in LOWZ and CMASS

We measure the galaxy-ellipticity (GI) correlations for the Sloan Digital Sky Survey Data Release 12 LOWZ and CMASS samples with the shape measurements from the DESI Legacy Imaging Surveys. We model the GI correlations in an N-body simulation with our recent accurate stellar–halo mass relation from the Photometric object Around Cosmic webs (PAC) method. The large data set and our accurate modeling turns out an accurate measurement of the alignment angle between central galaxies and their host halos. We find that the alignment of central elliptical galaxies with their host halos increases monotonically with galaxy stellar mass or host halo mass, which can be well described by a power law for the massive galaxies. We also find that central elliptical galaxies are more aligned with their host halos in LOWZ than in CMASS, which might indicate an evolution of galaxy–halo alignment, though future studies are needed to verify this is not induced by the sample selections. In contrast, central disk galaxies are aligned with their host halos about 10 times more weakly in the GI correlation. These results have important implications for intrinsic alignment (IA) correction in weak lensing studies, IA cosmology, and theory of massive galaxy formation.

79 ASTRONOMY AND ASTROPHYSICS↗

Visualization techniques to aid in the analysis of multi-spectral astrophysical data sets

The goal of this project was to support the scientific analysis of multi-spectral astrophysical data by means of scientific visualization. Scientific visualization offers its greatest value if it is not used as a method separate or alternative to other data analysis methods but rather in addition to these methods. Together with quantitative analysis of data, such as offered by statistical analysis, image or signal processing, visualization attempts to explore all information inherent in astrophysical data in the most effective way. Data visualization is one aspect of data analysis. Our taxonomy as developed in Section 2 includes identification and access to existing information, preprocessing and quantitative analysis of data, visual representation and the user interface as major components to the software environment of astrophysical data analysis. In pursuing our goal to provide methods and tools for scientific visualization of multi-spectral astrophysical data, we therefore looked at scientific data analysis as one whole process, adding visualization tools to an already existing environment and integrating the various components that define a scientific data analysis environment. As long as the software development process of each component is separate from all other components, users of data analysis software are constantly interrupted in their scientific work in order to convert from one data format to another, or to move from one storage medium to another, or to switch from one user interface to another. We also took an in-depth look at scientific visualization and its underlying concepts, current visualization systems, their contributions, and their shortcomings. The role of data visualization is to stimulate mental processes different from quantitative data analysis, such as the perception of spatial relationships or the discovery of patterns or anomalies while browsing through large data sets. Visualization often leads to an intuitive understanding of the meaning of data values and their relationships by sacrificing accuracy in interpreting the data values. In order to be accurate in the interpretation, data values need to be measured, computed on, and compared to theoretical or empirical models (quantitative analysis). If visualization software hampers quantitative analysis (which happens with some commercial visualization products), its use is greatly diminished for astrophysical data analysis. The software system STAR (Scientific Toolkit for Astrophysical Research) was developed as a prototype during the course of the project to better understand the pragmatic concerns raised in the project. STAR led to a better understanding on the importance of collaboration between astrophysicists and computer scientists.

Brugel, Edward W.↗

Visualization techniques to aid in the analysis of multispectral astrophysical data sets

The goal of this project was to support the scientific analysis of multi-spectral astrophysical data by means of scientific visualization. Scientific visualization offers its greatest value if it is not used as a method separate or alternative to other data analysis methods but rather in addition to these methods. Together with quantitative analysis of data, such as offered by statistical analysis, image or signal processing, visualization attempts to explore all information inherent in astrophysical data in the most effective way. Data visualization is one aspect of data analysis. Our taxonomy as developed in Section 2 includes identification and access to existing information, preprocessing and quantitative analysis of data, visual representation and the user interface as major components to the software environment of astrophysical data analysis. In pursuing our goal to provide methods and tools for scientific visualization of multi-spectral astrophysical data, we therefore looked at scientific data analysis as one whole process, adding visualization tools to an already existing environment and integrating the various components that define a scientific data analysis environment. As long as the software development process of each component is separate from all other components, users of data analysis software are constantly interrupted in their scientific work in order to convert from one data format to another, or to move from one storage medium to another, or to switch from one user interface to another. We also took an in-depth look at scientific visualization and its underlying concepts, current visualization systems, their contributions and their shortcomings. The role of data visualization is to stimulate mental processes different from quantitative data analysis, such as the perception of spatial relationships or the discovery of patterns or anomalies while browsing through large data sets. Visualization often leads to an intuitive understanding of the meaning of data values and their relationships by sacrificing accuracy in interpreting the data values. In order to be accurate in the interpretation, data values need to be measured, computed on, and compared to theoretical or empirical models (quantitative analysis). If visualization software hampers quantitative analysis (which happens with some commercial visualization products), its use is greatly diminished for astrophysical data analysis. The software system STAR (Scientific Toolkit for Astrophysical Research) was developed as a prototype during the course of the project to better understand the pragmatic concerns raised in the project. STAR led to a better understanding on the importance of collaboration between astrophysicists and computer scientists. Twenty-one examples of the use of visualization for astrophysical data are included with this report. Sixteen publications related to efforts performed during or initiated through work on this project are listed at the end of this report.

Brugel, E. W.↗

Oleaginous Yeast Biology Elucidated With Comparative Transcriptomics

ABSTRACT Extremophilic yeasts have favorable metabolic and tolerance traits for biomanufacturing‐ like lipid biosynthesis, flavinogenesis, and halotolerance – yet the connection between these favorable phenotypes and strain genotype is not well understood. To this end, this study compares the phenotypes and gene expression patterns of biotechnologically relevant yeasts Yarrowia lipolytica , Debaryomyces hansenii , and Debaryomyces subglobosus grown under nitrogen starvation, iron starvation, and salt stress. To analyze the large data set across species and conditions, two approaches were used: a “network‐first” approach where a generalized metabolic network serves as a scaffold for mapping genes and a “cluster‐first” approach where unsupervised machine learning co‐expression analysis clusters genes. Both approaches provide insight into strain behavior. The network‐first approach corroborates that Yarrowia upregulates lipid biosynthesis during nitrogen starvation and provides new evidence that riboflavin overproduction in Debaryomyces yeasts is overflow metabolism that is routed to flavin cofactor production under salt stress. The cluster‐first approach does not rely on annotation; therefore, the coexpression analysis can identify known and novel genes involved in stress responses, mainly transcription factors and transporters. Therefore, this work links the genotype to the phenotype of biotechnologically relevant yeasts and demonstrates the utility of complementary computational approaches to gain insight from transcriptomics data across species and conditions.

Weintraub, Sarah J. [Department of Bioinformatics ↗

Soil organic carbon is not just for soil scientists: measurement recommendations for diverse practitioners

Soil organic carbon (SOC) regulates terrestrial ecosystem functioning, provides diverse energy sources for soil microorganisms, governs soil structure, and regulates the availability of organically bound nutrients. Investigators in increasingly diverse disciplines recognize how quantifying SOC attributes can provide insight about ecological states and processes. Today, multiple research networks collect and provide SOC data, and robust, new technologies are available for managing, sharing, and analyzing large data sets. Here, we advocate that the scientific community capitalize on these developments to augment SOC data sets via standardized protocols. We describe why such efforts are important and the breadth of disciplines for which it will be helpful, and outline a tiered approach for standardized sampling of SOC and ancillary variables that ranges from simple to more complex. We target scientists ranging from those with little to no background in soil science to those with more soil-related expertise, and offer examples of the ways in which the resulting data can be organized, shared, and discoverable.

54 ENVIRONMENTAL SCIENCES↗

What's Left for a Computational Chemist To Do in the Age of Machine Learning?

Machine learning (ML) has become a central focus of the computational chemistry community. In this paper, I will first discuss my personal history in the field. Then I will provide a broader view of how this resurgence in ML interest echoes and advances upon earlier efforts. Although numerous changes have brought about this latest wave, one of the most significant is the increased accuracy and efficiency of low-cost methods (e.g., density functional theory or DFT) that have made it possible to generate large data sets for ML models. ML has also been used to bypass, guide, or improve DFT. The field of computational chemistry thus finds itself at a crossroads as ML both augments and supersedes traditional efforts. I will present what I believe the role of the computational chemist will be in this evolving landscape, with specific focus on my experience in the development of autonomous workflows in computational materials discovery for open-shell transition-metal chemistry.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗