Vicente Pena Perez 08739 Extreme-Scale Data Science & Analytics, Sandia National Laboratories.
Abstract not provided.
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Abstract not provided.
This presentation is for the paper INL/CON-22-65800
Final report
Data science has primarily focused on big data, but for many physics, chemistry, and engineering applications, data are often small, correlated and, thus, low dimensional, and sourced from both computations and experiments with various levels of noise. Typical statistics and machine learning methods do not work for these cases. Expert knowledge is essential, but a systematic framework for incorporating it into physics-based models under uncertainty is lacking. Here, we develop a mathematical and computational framework for probabilistic artificial intelligence (AI)–based predictive modeling combining data, expert knowledge, multiscale models, and information theory through uncertainty quantification and probabilistic graphical models (PGMs). We apply PGMs to chemistry specifically and develop predictive guarantees for PGMs generally. Our proposed framework, combining AI and uncertainty quantification, provides explainable results leading to correctable and, eventually, trustworthy models. The proposed framework is demonstrated on a microkinetic model of the oxygen reduction reaction.
Background: The COVID-19 pandemic remains a significant global threat. However, despite urgent need, there remains uncertainty surrounding best practices for pharmaceutical interventions to treat COVID-19. In particular, conflicting evidence has emerged surrounding the use of hydroxychloroquine and azithromycin, alone or in combination, for COVID-19. The COVID-19 Evidence Accelerator convened by the Reagan-Udall Foundation for the FDA, in collaboration with Friends of Cancer Research, assembled experts from the health systems research, regulatory science, data science, and epidemiology to participate in a large parallel analysis of different data sets to further explore the effectiveness of these treatments. Methods: Electronic health record (EHR) and claims data were extracted from seven separate databases. Parallel analyses were undertaken on data extracted from each source. Each analysis examined time to mortality in hospitalized patients treated with hydroxychloroquine, azithromycin, and the two in combination as compared to patients not treated with either drug. Cox proportional hazards models were used, and propensity score methods were undertaken to adjust for confounding. Frequencies of adverse events in each treatment group were also examined. Results: Neither hydroxychloroquine nor azithromycin, alone or in combination, were significantly associated with time to mortality among hospitalized COVID-19 patients. No treatment groups appeared to have an elevated risk of adverse events. Conclusion: Administration of hydroxychloroquine, azithromycin, and their combination appeared to have no effect on time to mortality in hospitalized COVID-19 patients. Continued research is needed to clarify best practices surrounding treatment of COVID-19.
Quantum machine learning—and specifically Variational Quantum Algorithms (VQAs)—offers a powerful, flexible paradigm for programming near-term quantum computers, with applications in chemistry, metrology, materials science, data science, and mathematics. Here, one trains an ansatz, in the form of a parameterized quantum circuit, to accomplish a task of interest. However, challenges have recently emerged suggesting that deep ansatzes are difficult to train, due to flat training landscapes caused by randomness or by hardware noise. This motivates our work, where we present a variable structure approach to build ansatzes for VQAs. Our approach, called VAns (Variable Ansatz), applies a set of rules to both grow and (crucially) remove quantum gates in an informed manner during the optimization. Consequently, VAns is ideally suited to mitigate trainability and noise-related issues by keeping the ansatz shallow. We employ VAns in the variational quantum eigensolver for condensed matter and quantum chemistry applications, in the quantum autoencoder for data compression and in unitary compilation problems showing successful results in all cases.
Data for the report (ORNL/TM-2025/3807) that describes the calculations and the Monte Carlo Ray Tracing simulations performed using the McStas package to determine the coatings and geometry for the NB-6 guide. It provides the information to inform the mechanical design, validation tests and verification that it meets the science requirements.
Prenyltransfer is an early-stage carbon-hydrogen bond (C-H) functionalization prevalent in the biosynthesis of a diverse array of biologically active bacterial, fungal, plant, and metazoan diketopiperazine (DKP) alkaloids. Toward the development of a unified strategy for biocatalytic construction of prenylated DKP indole alkaloids, we sought to identify and characterize a substrate-permissive C2 reverse prenyltransferase (PT). As the first tailoring event within the biosynthesis of cytotoxic notoamide metabolites, PT NotF catalyzes C2 reverse prenyltransfer of brevianamide F. Solving a crystal structure of NotF (in complex with native substrate and prenyl donor mimic dimethylallyl S-thiolodiphosphate (DMSPP)) revealed a large, solvent-exposed active site, intimating NotF may possess a significantly broad substrate scope. To assess the substrate selectivity of NotF, we synthesized a panel of 30 sterically and electronically differentiated tryptophanyl DKPs, the majority of which were selectively prenylated by NotF in synthetically useful conversions (2 to > 99%). Quantitative representation of this substrate library and development of a descriptive statistical model provided insight into the molecular origins of NotF's substrate promiscuity. This approach enabled the identification of key substrate descriptors (electrophilicity, size, and flexibility) that govern the rate of NotF-catalyzed prenyltransfer, and the development of an "induced fit docking (IFD)-guided" engineering strategy for improved turnover of our largest substrates. We further demonstrated the utility of NotF in tandem with oxidative cyclization using flavin monooxygenase, BvnB. This one-pot, in vitro biocatalytic cascade enabled the first chemoenzymatic synthesis of the marine fungal natural product, (-)-eurotiumin A, in three steps and 60% overall yield.
In 2019, highway congestion wasted over 3 billion gallons of fuel and caused 8.8 billion hours of lost productivity.1 Research has shown that introducing near-real time traffic controls can significantly reduce congestion. Validated and calibrated traffic simulations enable the modeling of transportation systems and the evaluation of different traffic control actions and schemes given a variety of circumstances that represent likely future scenarios. The developed scenarios can inform the deployment of controls in near real-time to improve freight and passenger vehicle congestion and energy use. In this work, we present simulations used to model the traffic in the Chattanooga, Tennessee, metropolitan area. Simulations were constructed and calibrated using a variety of local, data science enhanced, data sources utilizing open source software including the Simulation of Urban Mobility (SUMO) simulator. High-Performance Computing (HPC) provides a scalable platform with enough computing for the high-fidelity simulation of many scenarios and the application of advance data science especially for large-scale systems. Our simulations include microscopic simulations at a corridor level for traffic signal control, and mesoscopic simulations to evaluate regional operational controls and infrastructure.
Dislocations play a vital role in the mechanical behavior of crystalline materials during deformation. To capture dislocation phenomena across all relevant scales, a multiscale modeling framework of plasticity has emerged, with the goal of reaching a quantitative understanding of microstructure–property relations, for instance, to predict the strength and toughness of metals and alloys for engineering applications. This review describes the state of the art of the major dislocation modeling techniques, and then discusses how recent progress can be leveraged to advance the frontiers in simulations of dislocations. Furthermore, the frontiers of dislocation modeling include opportunities to establish quantitative connections between the scales, validate models against experiments, and use data science methods (e.g., machine learning) to gain an understanding of and enhance the current predictive capabilities.
To harness the potential of microbiome science across the broad range of relevant disciplines, new approaches to data infrastructure and transdisciplinary collaboration are necessary. The National Microbiome Data Collaborative (NMDC) is a new initiative to support microbiome data exploration and discovery through a collaborative integrative data science ecosystem.
Free use and redistribution of data (i.e., Open Data) increases the reproducibility, transparency, and pace of aquatic sciences research. However, barriers to both data users and data providers may limit the adoption of Open Data practices. Here, we describe common Open Data challenges faced by data users and data providers within the aquatic sciences community (i.e., oceanography, limnology, hydrology, and others). These challenges were synthesized from literature, authors’ experiences, and a broad survey of 174 data users and data providers across academia, government agencies, industry, and other sectors. Through this work, we identified seven main challenges: 1) metadata shortcomings, 2) variable data quality and reusability, 3) open data inaccessibility, 4) lack of standardization, 5) authorship and acknowledgement issues 6) lack of funding, and 7) unequal barriers around the globe. Our key recommendation is to improve resources to advance Open Data practices. This includes dedicated funds for capacity building, hiring and maintaining of skilled personnel, and robust digital infrastructures for preparation, storage, and long-term maintenance of Open Data. Further, to incentivize data sharing we reinforce the need for standardized best practices to handle data acknowledgement and citations for both data users and data providers. We also highlight and discuss regional disparities in resources and research practices within a global perspective.
The overarching objective of this effort has been to combine state-of-the-art data science techniques, first principles analyses, and molecular-level characterization of electrocatalyst structure and reactivity to identify both in-situ mechanisms for degradation and transformation of electrocatalysts with highly complex catalytic structures and the impact of these transformations on catalytic activity. The primary catalysts of interest have been multielemental alloys, including high entropy alloys (HEA’s), which are characterized by a high degree of disorder and up to 20 different elements within a single nanoparticle. We have applied these strategies primarily to energy-critical oxygen cycle electrocatalytic reactions, including oxygen reduction (ORR), but we have also considered extensions to non-electrochemical chemistries such as ammonia synthesis and decomposition. We have made strong progress in the development of computational methods on both the level of machine learning methods development as well as first principles-based treatments of HEA’s, and we have leveraged these insights to propose promising HEA catalysts for the ORR. On the experimental side, we developed new HEA synthesis and characterization protocols relevant to these reactions and developed a database combining our experimental results with corresponding computational tools.
Science Capsule captures the processing and data life cycle across machines, organizations, people, and science domains. Our approach combines user engagement methods from social sciences with innovations in workflow and data management.
In an era of increasingly sophisticated threats to critical infrastructure, cybersecurity professionals must be more than just aware; they must be immersed, agile, and equipped to operate in environments where failure is not an option. Nowhere is this truer than in the nuclear sector, where cyber-physical systems, regulatory scrutiny, and insider threat potential demand a new generation of hands-on, technically fluent defenders. This paper presents a unified training approach that integrates Cybersecurity Test and Evaluation (T&E) with data science techniques using Jupyter Notebooks as the interactive lab environment. The program centers on a modular, scenario-driven curriculum designed to build not just knowledge but practical capability in the assessment and defense of radiation detection systems, firmware interfaces, and operational security postures.
The benefits of addressing the water, energy, and food sectors in an integrated manner is gaining significant recognition. An integrated approach can provide improved resource use efficiencies, more coherent environmental policies, and an overall strategy for achieving sustainability in the three sectors, as outlined in the United Nations Sustainable Development Goals (SDG) 2 (Food), 6 (Water), and 7 (Energy). Societies are concerned with ensuring food security, avoiding wars over water, and creating opportunity by ensuring access to energy. To be effective, however, this approach needs to be adopted at all levels of societies including government, private and civil society and reinforced by management and planning methods. This special issue identifies different approaches that are either being conceptualized or tested to support the Water-Energy-Food (WEF) Nexus approach. The articles contribute to answering the question, “Is achieving Water-Energy-Food (WEF) Nexus Sustainability a science and data need or an integrated public policy need?” In either case, both natural and social sciences will need to combine to support science issues or integrated policy issues. The papers explore the ways in which science, data, and policy development could help to define integrative principles and policies for the three sectors. This approach could expand beyond the water, energy, and food sectors to include health, environment, trade, commerce, and international assistance thereby providing broad support to the SDGs. Moreover, this issue demonstrates that data combined with new technologies (tools and models) can support better decision-making when adopted by governments and the sectors.
Scientific collaboration is a long-standing subject of CSCW scholarship that typically focuses on the development and use of computing systems to facilitate research. The research presented in this article investigates the sociality of science by identifying and describing particular, common forms of organizing that researchers in four different scientific realms employ to conduct work in both local contexts and as part of distributed, global projects. This paper introduces five prototypical forms of organizing we categorize as coordinative entities: the Principal Group, Intermittent Exchange, Sustained Aggregation, Federation, and Facility Organization. Coordinative entities as a categorization help specify, articulate, compare, and trace overlapping and evolving arrangements scientists use to facilitate data intensive research. We use this typology to unpack complexities of data intensive scientific collaboration in four cases, showing how scientists invoke different coordinative entities across three types of research activities: data collection, processing, and analysis. Finally, our contribution scrutinizes the sociality of scientific work to illustrate how these actors engage in relational work within and among diverse, dispersed forms of organizing across project, funding, and disciplinary boundaries.
The Transformational Challenge Reactor is being designed at Oak Ridge National Laboratory to demonstrate the feasibility of constructing a reactor core using advanced manufacturing technology. This technology includes additive manufacturing combined with machine learning, materials science, and data science technologies in an effort to facilitate the expansion of additive manufacturing into advanced nuclear energy systems and other applications requiring a high level of quality assurance. The Transformational Challenge Reactor is employing additive manufacturing and artificial intelligence to deliver a new approach. Beginning in FY21, the focus of the program has shifted away from demonstrating a reactor, and instead, towards delivering on four key thrust areas: (1) artificial intelligence-informed design, (2) advanced materials, (3) integrated sensing and control, and (4) the digital platform. Of these four thrust areas, the most pertinent to this report is the digital platform. The digital platform has the potential to be a key enabler for a paradigm shift in how components, those derived from advanced manufacturing technologies, are certified for use in nuclear applications. This is achieved primarily using machine learning to discover correlations from the abundance of data produced through additive manufacturing and those physical properties critical to the performance of the component.