Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data sciences”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Data for NB6 HBRR Science Design ORNL/TM-2025/3807

Data for the report (ORNL/TM-2025/3807) that describes the calculations and the Monte Carlo Ray Tracing simulations performed using the McStas package to determine the coatings and geometry for the NB-6 guide. It provides the information to inform the mechanical design, validation tests and verification that it meets the science requirements.

47 OTHER INSTRUMENTATION↗

Data Science-Driven Analysis of Substrate-Permissive Diketopiperazine Reverse Prenyltransferase NotF: Applications in Protein Engineering and Cascade Biocatalytic Synthesis of (-)-Eurotiumin A

Prenyltransfer is an early-stage carbon-hydrogen bond (C-H) functionalization prevalent in the biosynthesis of a diverse array of biologically active bacterial, fungal, plant, and metazoan diketopiperazine (DKP) alkaloids. Toward the development of a unified strategy for biocatalytic construction of prenylated DKP indole alkaloids, we sought to identify and characterize a substrate-permissive C2 reverse prenyltransferase (PT). As the first tailoring event within the biosynthesis of cytotoxic notoamide metabolites, PT NotF catalyzes C2 reverse prenyltransfer of brevianamide F. Solving a crystal structure of NotF (in complex with native substrate and prenyl donor mimic dimethylallyl S-thiolodiphosphate (DMSPP)) revealed a large, solvent-exposed active site, intimating NotF may possess a significantly broad substrate scope. To assess the substrate selectivity of NotF, we synthesized a panel of 30 sterically and electronically differentiated tryptophanyl DKPs, the majority of which were selectively prenylated by NotF in synthetically useful conversions (2 to > 99%). Quantitative representation of this substrate library and development of a descriptive statistical model provided insight into the molecular origins of NotF's substrate promiscuity. This approach enabled the identification of key substrate descriptors (electrophilicity, size, and flexibility) that govern the rate of NotF-catalyzed prenyltransfer, and the development of an "induced fit docking (IFD)-guided" engineering strategy for improved turnover of our largest substrates. We further demonstrated the utility of NotF in tandem with oxidative cyclization using flavin monooxygenase, BvnB. This one-pot, in vitro biocatalytic cascade enabled the first chemoenzymatic synthesis of the marine fungal natural product, (-)-eurotiumin A, in three steps and 60% overall yield.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Challenges of open data in aquatic sciences: issues faced by data users and data providers

Free use and redistribution of data (i.e., Open Data) increases the reproducibility, transparency, and pace of aquatic sciences research. However, barriers to both data users and data providers may limit the adoption of Open Data practices. Here, we describe common Open Data challenges faced by data users and data providers within the aquatic sciences community (i.e., oceanography, limnology, hydrology, and others). These challenges were synthesized from literature, authors’ experiences, and a broad survey of 174 data users and data providers across academia, government agencies, industry, and other sectors. Through this work, we identified seven main challenges: 1) metadata shortcomings, 2) variable data quality and reusability, 3) open data inaccessibility, 4) lack of standardization, 5) authorship and acknowledgement issues 6) lack of funding, and 7) unequal barriers around the globe. Our key recommendation is to improve resources to advance Open Data practices. This includes dedicated funds for capacity building, hiring and maintaining of skilled personnel, and robust digital infrastructures for preparation, storage, and long-term maintenance of Open Data. Further, to incentivize data sharing we reinforce the need for standardized best practices to handle data acknowledgement and citations for both data users and data providers. We also highlight and discuss regional disparities in resources and research practices within a global perspective.

54 ENVIRONMENTAL SCIENCES↗

Data Science-Driven Discovery of Multimetallic Oxygen-cycle Electrocatalysts for Enhanced Energy Conversion

The overarching objective of this effort has been to combine state-of-the-art data science techniques, first principles analyses, and molecular-level characterization of electrocatalyst structure and reactivity to identify both in-situ mechanisms for degradation and transformation of electrocatalysts with highly complex catalytic structures and the impact of these transformations on catalytic activity. The primary catalysts of interest have been multielemental alloys, including high entropy alloys (HEA’s), which are characterized by a high degree of disorder and up to 20 different elements within a single nanoparticle. We have applied these strategies primarily to energy-critical oxygen cycle electrocatalytic reactions, including oxygen reduction (ORR), but we have also considered extensions to non-electrochemical chemistries such as ammonia synthesis and decomposition. We have made strong progress in the development of computational methods on both the level of machine learning methods development as well as first principles-based treatments of HEA’s, and we have leveraged these insights to propose promising HEA catalysts for the ORR. On the experimental side, we developed new HEA synthesis and characterization protocols relevant to these reactions and developed a database combining our experimental results with corresponding computational tools.

36 MATERIALS SCIENCE↗

Hands-On, Heads-Up: Blending Cyber T&E with Data Science-Driven Training in Jupyter Notebooks

In an era of increasingly sophisticated threats to critical infrastructure, cybersecurity professionals must be more than just aware; they must be immersed, agile, and equipped to operate in environments where failure is not an option. Nowhere is this truer than in the nuclear sector, where cyber-physical systems, regulatory scrutiny, and insider threat potential demand a new generation of hands-on, technically fluent defenders. This paper presents a unified training approach that integrates Cybersecurity Test and Evaluation (T&E) with data science techniques using Jupyter Notebooks as the interactive lab environment. The program centers on a modular, scenario-driven curriculum designed to build not just knowledge but practical capability in the assessment and defense of radiation detection systems, firmware interfaces, and operational security postures.

98 - NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL↗

Editorial: Achieving Water-Energy-Food Nexus Sustainability: A Science and Data Need or a Need for Integrated Public Policy?

The benefits of addressing the water, energy, and food sectors in an integrated manner is gaining significant recognition. An integrated approach can provide improved resource use efficiencies, more coherent environmental policies, and an overall strategy for achieving sustainability in the three sectors, as outlined in the United Nations Sustainable Development Goals (SDG) 2 (Food), 6 (Water), and 7 (Energy). Societies are concerned with ensuring food security, avoiding wars over water, and creating opportunity by ensuring access to energy. To be effective, however, this approach needs to be adopted at all levels of societies including government, private and civil society and reinforced by management and planning methods. This special issue identifies different approaches that are either being conceptualized or tested to support the Water-Energy-Food (WEF) Nexus approach. The articles contribute to answering the question, “Is achieving Water-Energy-Food (WEF) Nexus Sustainability a science and data need or an integrated public policy need?” In either case, both natural and social sciences will need to combine to support science issues or integrated policy issues. The papers explore the ways in which science, data, and policy development could help to define integrative principles and policies for the three sectors. This approach could expand beyond the water, energy, and food sectors to include health, environment, trade, commerce, and international assistance thereby providing broad support to the SDGs. Moreover, this issue demonstrates that data combined with new technologies (tools and models) can support better decision-making when adopted by governments and the sectors.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Digital Platform Informed Certification of Components Derived from Advanced Manufacturing Technologies

The Transformational Challenge Reactor is being designed at Oak Ridge National Laboratory to demonstrate the feasibility of constructing a reactor core using advanced manufacturing technology. This technology includes additive manufacturing combined with machine learning, materials science, and data science technologies in an effort to facilitate the expansion of additive manufacturing into advanced nuclear energy systems and other applications requiring a high level of quality assurance. The Transformational Challenge Reactor is employing additive manufacturing and artificial intelligence to deliver a new approach. Beginning in FY21, the focus of the program has shifted away from demonstrating a reactor, and instead, towards delivering on four key thrust areas: (1) artificial intelligence-informed design, (2) advanced materials, (3) integrated sensing and control, and (4) the digital platform. Of these four thrust areas, the most pertinent to this report is the digital platform. The digital platform has the potential to be a key enabler for a paradigm shift in how components, those derived from advanced manufacturing technologies, are certified for use in nuclear applications. This is achieved primarily using machine learning to discover correlations from the abundance of data produced through additive manufacturing and those physical properties critical to the performance of the component.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Data Driven User Emulator

This software generates realistic network traffic to test intrusion detection systems. The most realistic traffic is generated when software is used to drive actual applications, thereby behaving like a real user. Tools that generate such realistic user behavior are called user emulators. However, no existing user emulators use models based on data science and real user data, and they suffer from a decrease in the fidelity of the generated traffic. We develop a user emulator that uses real user data and data science to generate higher fidelity emulation and increase the accuracy of our experimental results.

Oesch, TimothyS.↗

Sandia Academic Alliance Program Collaboration Report: 2020-2021 Accomplishments

University partnerships play an essential role in sustaining Sandia’s vitality as a national laboratory. The SAA is an element of Sandia’s broader University Partnerships program, which facilitates recruiting and research collaborations with dozens of universities annually. The SAA program has two three-year goals. SAA aims to realize a step increase in hiring results, by growing the total annual inexperienced hires from each out-of-state SAA university. SAA also strives to establish and sustain strategic research partnerships by establishing several federally sponsored collaborations and multi-institutional consortiums in science & technology (S&T) priorities such as autonomy, advanced computing, hypersonics, quantum information science, and data science. The SAA program facilitates access to talent, ideas, and Research & Development facilities through strong university partnerships. Earlier this year, the SAA program and campus executives hosted John Myers, Sandia’s former Senior Director of Human Resources (HR) and Communications, and senior-level staff at Georgia Tech, U of Illinois, Purdue, UNM, and UT Austin. These campus visits provided an opportunity to share the history of the partnerships from the university leadership, tours of research facilities, and discussions of ongoing technical work and potential recruiting opportunities. These visits also provided valuable feedback to HR management that will help Sandia realize a step increase in hiring from SAA schools. The 2020-2021 Collaboration Report is a compilation of accomplishments in 2020 and 2021 from SAA and Sandia’s valued SAA university partners.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

DevOps and Data: Faster-Time-to-Knowledge through SageOps, MLOps, and DataOps

This report describes the approach, investigation, and prototyping efforts to develop an efficient, reusable methodology and reference framework for applying DevOps to disparate data for data science and data analytics at scale, based on focused application of this methodology and reusable reference framework within Sandia National Laboratories’ Pulsed Power community. Additionally, this report reviews: engineered instantiation of the reference framework used for development and production solutions, our experiences and results in using the reference framework, and future plans regarding research and development.

97 MATHEMATICS AND COMPUTING↗

An Indicator-based Approach to Sustainable Management of Natural Resources (Chapter 12)

Assessing the sustainability of natural resource management choices for agricultural and forest lands requires quantification of potential changes to a set of environmental and socioeconomic indicators selected to characterize reference scenarios relative to projected future scenarios. Correctly framing the questions with local stakeholders is a critical first step in the sustainability assessment, and the questions that can be addressed are often limited by data availability. Selecting and prioritizing indicators with stakeholders to address their needs and concerns improves the likelihood of investment in monitoring and evaluation of those indicators over time. Computational techniques for analyzing interactions between the selected indicators are inherently affected by the scales and formats of the assembled indicator datasets. Data analytics have the potential to improve understanding of the potential synergies and tradeoffs involved with meeting multiple environmental and socioeconomic goals simultaneously, but timely and appropriate indicator datasets are not always available—even in this new era of “big data.” Continued improvements in data science and data analytics are needed to broaden understanding and acceptance of problems and to provide valuable information for natural resource management. Advances in these areas will enable society to design future landscapes that meet multiple objectives, including the provisioning of agricultural and forest resources along with a variety of ecosystem services (e.g., clean water and healthy soils).

Parish, Esther↗

Tutorial: Lessons Learned for Behavior Analysts from Data Scientists

Big data is a computing term used to refer to large and complex data sets, typically consisting of terabytes or more of diverse data that is produced rapidly. The analysis of such complex data sets requires advanced analysis techniques with the capacity to identify patterns and abstract meanings from the vast data. The field of data science combines computer science with mathematics/statistics and leverages artificial intelligence, in particular machine learning, to analyze big data. This field holds great promise for behavior analysis, where both clinical and research studies produce large volumes of diverse data at a rapid pace (i.e., big data). This article presents basic lessons for the behavior analytic researchers and clinicians regarding integration of data science into the field of behavior analysis. We provide guidance on how to collect, protect, and process the data, while highlighting the importance of collaborating with data scientists to select a proper machine learning model that aligns with the project goals and develop models with input from human experts. Here, we hope this serves as a guide to support the behavior analysts interested in the field of data science to advance their practice or research, and helps them avoid some common pitfalls.

42 ENGINEERING↗

2022 American Conference on Neutron Scattering (ACNS 2022)

The 11th American Conference on Neutron Scattering (ACNS 2022) will be held on June 5-9, 2022, in Boulder, CO. The Conference will provide essential information on the breadth and depth of current neutron-related research worldwide. Hosted by the Neutron Scattering Society of America, this year’s Conference will feature a combination of invited and contributed talks, poster sessions, and tutorials. Topics of the conference are: Advances in Neutron Facilities, Instrumentation and Software: Developments in sources, instrumentation, sample environments and control software. Hard Condensed Matter: Magnetism, correlated metals, quantum/topological materials, superconductors, ferroelectrics, multiferroics, glasses, and disorder phenomena. Submissions outlining examples of neutron scattering in industrial and engineering applications involving hard condensed matter systems are also encouraged. Soft Matter: Neutron studies of soft materials and related fields including in situ and in operando studies. Polymers, surfactants, emulsions, gels, nanoparticles, colloidal suspensions and more. Submissions of computational studies or applications of machine learning beneficial to neutron scattering experiments, as well as examples of neutron scattering in industrial and engineering applications are strongly encouraged. Biology, Biophysics and Biotechnology: Neutron studies of biological and biologically relevant systems. Proteins, bio membranes, biological assemblies, natural materials, nucleic acids, drug-delivery platforms and biomedical systems. Submissions of computational studies or applications of machine learning beneficial to biological neutron scattering experiments, as well as examples of neutron scattering in applied research involving biological systems, are strongly encouraged. Materials Chemistry and Energy: Neutron-based studies of functional materials and materials for energy applications. Examples include porous materials such as metal organic frameworks (MOFs), zeolites; phosphors; novel pigments; electrolytes; catalysts; ionic conductors/cathode materials; photovoltaic materials (hybrid perovskites); thermoelectrics; magnetocalorics/electrocalorics. Structural Materials and Engineering: Neutron scattering studies of materials and engineering processes including structural materials, concrete and metals, as well as engineering processes including combustion, corrosion, additive manufacturing, and others. Neutron Physics: Fundamental physical studies of the neutron and related areas. Emerging Applications in Neutron Scattering: Machine Learning and Data Science: Advances in computing power have contributed to rapidly evolving machine learning and data science fields that can be leveraged to the benefit of the neutron scattering community. The purpose of this session is to highlight recent advances in machine learning and data science and to serve as the foundation of a parallel data and computation track highlighting computation advances and applications in neutron scattering throughout the conference.

36 MATERIALS SCIENCE↗

47 Tuc in Rubin Data Preview 1. Exploring Early LSST Data and Science Potential

We present analyses of the early data from Rubin Observatory’s Data Preview 1 (DP1) for the field of the globular cluster 47 Tuc. The DP1 data set for 47 Tuc includes four nights of observations from the Rubin Commissioning Camera (LSSTComCam), covering multiple bands (ugriy). We address challenges of crowding in the inner region of the cluster and toward the SMC in DP1, and demonstrate improved star–galaxy separation by fitting fifth-degree polynomials to the stellar loci in color–color diagrams and applying multidimensional sigma clipping. We compile a catalog of 3576 probable 47 Tuc member stars selected via a combination of isochrone, Gaia proper-motion, and color–color space matched filtering. We explore the sources of photometric scatter in the 47 Tuc color–color sequence, evaluating contributions from various potential sources, including differential extinction within the cluster. Finally, of the 72 well-characterized variables in the field, we recover three known variable stars, including two RR Lyrae and one eclipsing binary, in the coadd-based object catalog, and identify 62 in the difference image-based object catalog. Although the DP1 lightcurves have sparse temporal sampling, they appear to follow the patterns of densely sampled literature lightcurves well. Despite some data limitations for crowded-field stellar analysis, DP1 demonstrates the promising scientific potential for future LSST data releases.

Choi, Yumi [NSF National Optical-Infrared Astronom↗

A proximal trust-region method for nonsmooth optimization with inexact function and gradient evaluations

Many applications require minimizing the sum of smooth and nonsmooth functions. For example, basis pursuit denoising problems in data science require minimizing a measure of data misfit plus an $\ell^1$-regularizer. Similar problems arise in the optimal control of partial differential equations (PDEs) when sparsity of the control is desired. Here, we develop a novel trust-region method to minimize the sum of a smooth nonconvex function and a nonsmooth convex function. Our method is unique in that it permits and systematically controls the use of inexact objective function and derivative evaluations. When using a quadratic Taylor model for the trust-region subproblem, our algorithm is an inexact, matrix-free proximal Newton-type method that permits indefinite Hessians. We prove global convergence of our method in Hilbert space and demonstrate its efficacy on three examples from data science and PDE-constrained optimization.

97 MATHEMATICS AND COMPUTING↗

LDM-151: Data Management Science Pipelines Design

The LSST Science Requirements Document (the LSST SRD) specifies a set of data product guidelines, designed to support science goals envisioned to be enabled by the LSST observing program. Following these guidelines, the details of these data products have been described in the LSST Data Products Definition Document (DPDD), and captured in a formal flow-down from the SRD via the LSST System Requirements (LSR), Observatory System Specifications (OSS), to the Data Management System Requirements (DMSR). The LSST Data Management subsystem's responsibilities include the design, implementation, deployment and execution of software pipelines necessary to generate these data products. This document describes the design of the scientific aspects of those pipelines.

79 ASTRONOMY AND ASTROPHYSICS↗

DSI Python API Demo May 2023 [Slides]

Data Science Infrastructure (DSI) is developing searchable databases and workflows for simulation and experimental data derived from ASC clients. Short term goal: Make data easily accessible through metadata indexing and querying while respecting data permissions. Longer term goal: Use this data for data science activities.

97 MATHEMATICS AND COMPUTING↗

Bioinformatic teaching resources - for educators, by educators - using KBase, a free, user-friendly, open source platform

Over the past year, biology educators and staff at the Department of Energy Systems Biology Knowledgebase (KBase) initiated a collaborative effort to develop a curriculum for bioinformatics education. KBase is a free and easily accessible data science platform that integrates many bioinformatics resources into a graphical user interface built upon reproducible analysis notebooks. KBase held conversations with college and high school instructors to understand how KBase could potentially support their educational goals. These conversations morphed into a working group of biological and data science instructors that adapted the KBase platform to their curriculum needs, specifically around concepts in Genomics, Metagenomics, Pangenomics, and Phylogenetics. The KBase Educators Working Group developed modular, adaptable, and customizable instructional units. Each instructional module contains teaching resources, publicly available data, analysis tools, and markdown capability to tailor instructions and learning goals for each class. The online user interface enables students to conduct hands-on data science research and analyses without requiring programming skills or their own computational resources (these are provided by KBase). Alongside these resources, KBase continues to work with instructors, supporting the development of additional curriculum modules. For anyone new to the platform, KBase, and the growing KBase Educators Organization, provides a community network, accompanied by community-sourced guidelines, instructional templates, and peer support to use KBase within a classroom whether virtual or in-person.

59 BASIC BIOLOGICAL SCIENCES↗