Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “SCIENCE”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Commutative Algebra Modeling in Materials Science – A Case Study on Metal–Organic Frameworks (MOFs)

Metal-organic frameworks (MOFs) are a class of important crystalline and highly porous materials whose hierarchical geometry and chemistry hinder interpretable predictions in materials properties. Commutative algebra is a branch of abstract algebra that has been rarely applied in data and material sciences. We introduce the first ever commutative algebra modeling and prediction in materials science. Specifically, category-specific commutative algebra (CSCA) is proposed as a new framework for MOF representation and learning. It integrates element-based categorization with multiscale algebraic invariants to encode both local coordination motifs and global network organization of MOFs. These algebraically consistent, chemically aware representations enable compact, interpretable, and data efficient modeling of MOF properties such as Henry’s constants and uptake capacities for common gases. Compared to traditional geometric and graph-based approaches, CSCA achieves comparable or superior predictive accuracy while substantially improving interpretability and stability across data sets. By aligning commutative algebra with the chemical hierarchy, the CSCA establishes a rigorous and generalizable paradigm for understanding structure and property relationships in porous materials and provides a nonlinear algebra-based framework for data-driven material discovery.

Khaemba, Caleb S.↗

Increasing the presence of BIPOC researchers in computational science

Here, Nature Computational Science asked a group of scientists to discuss strategies for increasing the presence of Black, Indigenous, People of Color (BIPOC) researchers in computational science, as well as the various considerations to be made for improving education and methods design.

99 GENERAL AND MISCELLANEOUS↗

The ePIC Simulation Campaign Workflow on the Open Science Grid

The ePIC collaboration is realizing the first experiment of the future Electron-Ion Collider (EIC) at the Brookhaven National Laboratory that will allow for a precision study of the nucleons and the nucleus at the scale of sea quarks and gluons through the study of electron-proton/ion collisions. This paper will discuss the current workflow for running centralized simulation campaigns for ePIC on the Open Science Grid (OSG) infrastructure. This involves monthly releases of ePIC software and container deployments to CVMFS, generation of input datasets in HepMC format according to collaboration-defined policy, using Snakemake in CI/CD for validation and benchmarking, and submitting jobs to the OSG condor scheduler for opportunistic running on available resources. File transfers utilize XrootD, and Rucio is used for data management. The workflow is continuously refined to improve daily throughput (currently 50-100k core hours per day) and minimize job failures. Since May 2023, monthly simulation campaigns employing the workflow have cumulatively used over 20 million core hours on the OSG and produced over 350 TB of simulation data. The campaigns incorporate simulations for the broad science program of the EIC and are actively used for the detector and physics studies in preparation of the Technical Design Report (TDR).

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

A customizable data management framework for high-repetition-rate high-energy-density science

The high-energy-density (HED) physics community is moving toward a new paradigm of high-repetition-rate (HRR) operation. To fully leverage the scientific power of HRR HED facilities, all of the components of each subsystem (laser, targetry, and performance diagnostics) must be connected and synchronized in a reliable and robust manner while the data acquired are tagged and archived in real time. To this end, GA has begun developing a generalized NoSQL-database framework, the MongoDB repository for information and archiving. An organizational strategy has been developed that shifts HED data organization from a shot-based to a diagnostic-based approach in order to increase archival and retrieval efficiency that lends itself to optimization applications. This work is a first step in pushing HRR HED science toward data management solutions that emphasize machine actionability and aim to stimulate community engagement to define data standards in HED science.

Instruments & Instrumentation↗

Data Science Shows that Entropy Correlates with Accelerated Zeolite Crystallization in Monte Carlo Simulations

We have performed a data science study of Monte Carlo simulation trajectories to understand factors that can accelerate formation of zeolite nanoporous crystals, a process that can take days or even weeks. In previous work, Monte Carlo simulations predicted and experiments confirmed that using a secondary organic structure-directing agent (OSDA) accelerates crystallization of all-silica LTA zeolite, with experiments finding a three-fold speedup [PCCP 24, 142-148 (2022)]. However, it remains unclear what physical factors cause the speed-up. Here, we apply data science to analyze the simulation trajectories to discover what drives accelerated zeolite crystallization in Monte Carlo going from a one-OSDA synthesis (1OSDA) to a two-OSDA version (2OSDA). We encoded simulation snapshots using the Smooth Overlap of Atomic Positions approach, which represents all 2- and 3-body correlations within a given cutoff distance. Principal component analyses failed to discriminate datasets of structures from 1OSDA and 2OSDA simulations, while the Support Vector Machine (SVM) approach succeeded at classifying such structures with an area-under-curve (AUC) score of 0.99 (where AUC = 1 is a perfect classification) with all 3-body correlations, and as high as 0.94 with only 2-body correlations. SVM decision functions reveal relatively broad / narrow histograms for 1OSDA / 2OSDA datasets, suggesting that the two simulations differ strongly in information heterogeneity. Informed by these results, we performed pair (2-body) entropy calculations during crystallization, resulting in entropy differences that semi-quantitatively account for the speedup observed in the previous Monte Carlo simulations. We conclude that altering synthesis conditions in ways that substantially changes the entropy of labile silica networks may accelerate zeolite crystallization, and we discuss possible approaches for achieving such acceleration.

77 NANOSCIENCE AND NANOTECHNOLOGY↗

Deep learning models map rapid plant species changes from citizen science and remote sensing data

Anthropogenic habitat destruction and climate change are reshaping the geographic distribution of plants worldwide. However, we are still unable to map species shifts at high spatial, temporal, and taxonomic resolution. Here, we develop a deep learning model trained using remote sensing images from California paired with half a million citizen science observations that can map the distribution of over 2,000 plant species. Our model— Deepbiosphere— not only outperforms many common species distribution modeling approaches (AUC 0.95 vs. 0.88) but can map species at up to a few meters resolution and finely delineate plant communities with high accuracy, including the pristine and clear-cut forests of Redwood National Park. These fine-scale predictions can further be used to map the intensity of habitat fragmentation and sharp ecosystem transitions across human-altered landscapes. In addition, from frequent collections of remote sensing data, Deepbiosphere can detect the rapid effects of severe wildfire on plant community composition across a 2-y time period. These findings demonstrate that integrating public earth observations and citizen science with deep learning can pave the way toward automated systems for monitoring biodiversity change in real-time worldwide.

Gillespie, Lauren E.↗

Opportunities for gas-phase science at short-wavelength free-electron lasers with undulator-based polarization control

Free-electron lasers (FELs) are the world's most brilliant light sources with rapidly evolving technological capabilities in terms of ultrabright and ultrashort pulses over a large range of photon energies. Their revolutionary and innovative developments have opened new fields of science regarding nonlinear light-matter interaction, the investigation of ultrafast processes from specific observer sites, and approaches to imaging matter with atomic resolution. A core aspect of FEL science is the study of isolated and prototypical systems in the gas phase with the possibility of addressing well-defined electronic transitions or particular atomic sites in molecules. Notably for polarization-controlled short-wavelength FELs, the gas phase offers new avenues for investigations of nonlinear and ultrafast phenomena in spin-orientated systems, for decoding the function of the chiral building blocks of life as well as steering reactions and particle emission dynamics in otherwise inaccessible ways. This roadmap comprises descriptions of technological capabilities of facilities worldwide, innovative diagnostics and instrumentation, as well as recent scientific highlights, novel methodology, and mathematical modeling. The experimental and theoretical landscape of using polarization controllable FELs for dichroic light-matter interaction in the gas phase will be discussed and comprehensively outlined to stimulate and strengthen global collaborative efforts of all disciplines. Published by the American Physical Society 2025

Ilchen, Markus↗

Machine Learning to Select Experiments Driven by Fundamental Science and Applications for Targeted Nuclear Data Improvement

This work describes a blueprint for a process that accelerates progress in science by quantitatively answering the following question: What is the optimal combination of fundamental-science and application-driven experiments to maximally reduce pertinent data uncertainties? Answering this question entails solving a high-dimensional and complex optimization problem that is best solved with advanced statistic techniques often classified as machine learning. We apply this process within the framework of nuclear data with the aim to select an experiment combination that will reduce uncertainties in 239 Pu nuclear data for neutron energies between 1 and 600 keV. In this field, fundamental-physics driven data, called differential, look at one nuclear physics observable at a time. They are contrasted to application-driven, integral, data where one or few resulting values inform a broad set of nuclear data across several nuclides and energies. The candidates for integral experiments are criticality measurements that were refined by a genetic algorithm to be maximally sensitive to 239 Pu fission cross sections in the desired energy range. Twenty-three candidate differential experiments were investigated and span multiple nuclear physics observables (e.g., total, capture cross sections) for isotopes appearing in the integral experiments. The optimal combination among these candidate experiments was investigated via generalized least squares fitting, augmented with Gaussian processes to ameliorate statistical irregularities in data, and the D-optimality criterion. The latter evaluates for each pair of candidates the joint reduction in uncertainties of all 12200 nuclear data appearing in the integral experiments compared to the knowledge we have from 168 past experiments, theory, and nuclear data. We chose as differential measurements those that investigate 63 Cu and 239 Pu total cross sections, based on D-optimality rank and feasibility constraints. Two integral (criticality) experiments were selected: An experiment with Al 2 ⁢O 3 and graphite interleaved with Pu and a thick Cu reflector explores 1–30 keV, while we target the 30–600 keV range with an experiment that swaps boron in place of graphite with a different geometry.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Privacy-Preserving Federated Learning for Science: Challenges and Research Directions

This paper discusses the key challenges and future research directions for privacy-preserving federated learning (PPFL), with a focus on its application to large-scale scientific AI models, in particular, foundation models~(FMs). PPFL enables collaborative model training across distributed datasets while preserving privacy-- an important collaborative approach for science. We discuss the need for efficient and scalable algorithms to address the increasing complexity of FMs, particularly when dealing with heterogeneous clients. In addition, we underscore the need for developing advance privacy-preserving techniques, such as differential privacy, to balance privacy and utility in large FMs emphasizing fairness and incentive mechanisms to ensure equitable participation among heterogeneous clients. Finally, we emphasize the need for a robust software stack supporting scalable and secure PPFL deployments across multiple high-performance computing facilities. We envision that PPFL would play a crucial role to advance scientific discovery and enable large-scale, privacy-aware collaborations across science domains.

Kim, Kibaek [Argonne National Laboratory (ANL)]↗

Research Software Engineering: Introducing a New Computing in Science & Engineering Department

Here, this article introduces the new Research Software Engineering (RSEng) department at Computing in Science & Engineering. Through a conversation with the department coeditors, we highlight why RSEng matters, how it differs from industrial software engineering, what it means to be an RSE, and the scholarly and practical questions that lie ahead. Along the way, we draw on emerging literature, case studies, and community perspectives to frame the profession and practice of RSEng within computational science and engineering.

Lamprecht, Anna-Lena [Univ. of Potsdam (Germany)] ↗

Material Needs and Measurement Challenges for Advanced Semiconductor Packaging: Understanding the Soft Side of Science

This Perspective builds upon insights from the National Institute of Standards and Technology (NIST)-organized workshop, “Materials and Metrology Needs for Advanced Semiconductor Packaging Strategies,” held at the 35th annual Electronics Packaging Symposium in Binghamton, NY, on September 5, 2024. It outlines critical challenges and opportunities related to polymer-based “soft” materials in advanced semiconductor packaging, with emphasis on polymer science, measurement science (metrology), and the strategic development of Research-Grade Test Materials (RGTMs). These efforts, led by the NIST CHIPS team, aim to advance the fundamental understanding of structure-property-processing relationships, promote standardized guidelines and innovative methods for material characterization, and accelerate the development, qualification, and adoption of next-generation packaging materials. The Perspective also distills key insights from the panel discussion with industry experts, emphasizing the need for close collaboration among materials scientists, process engineers, and metrology experts to enable a holistic strategy, further highlighting the importance of cross-sector partnerships among industry, academia, and government to address pressing challenges in packaging materials and processes.

97 MATHEMATICS AND COMPUTING↗

Machine learning-driven predictive resource management in complex science workflows

Here, the collaborative efforts of large communities in science experiments, often comprising thousands of global members, reflect a monumental commitment to exploration and discovery. Recently, advanced and complex data processing has gained increasing importance in science experiments. Data processing workflows typically consist of multiple intricate steps, and the precise specification of resource requirements is crucial for each step to allocate optimal resources for effective processing. Estimating resource requirements in advance is challenging due to a wide range of analysis scenarios, varying skill levels among community members, and the continuously increasing spectrum of computing options. One practical approach to mitigate these challenges involves initially processing a subset of each step to measure precise resource utilization from actual processing profiles before completing the entire step. While this two-staged approach enables processing on optimal resources for most of the workflow, it has drawbacks such as initial inaccuracies leading to potential failures and suboptimal resource usage, along with overhead from waiting for initial processing completion, which is critical for fast-turnaround analyses. In this context, our study introduces a novel pipeline of machine learning models within a comprehensive workflow management system, the Production and Distributed Analysis (PanDA) system. These models employ advanced machine learning techniques to predict key resource requirements, overcoming challenges posed by limited upfront knowledge of characteristics at each step. Accurate forecasts of resource requirements enable informed and proactive decision-making in workflow management, enhancing the efficiency of handling diverse, complex workflows across heterogeneous resources.

97 MATHEMATICS AND COMPUTING↗

FAIR Ecosystems for Science at Scale

High Performance Computing (HPC) centers provide resources to users who require greater scale to “get science done”. They deploy infrastructure with singular hardware architectures, cutting-edge software environments, and stricter security measures as compared with users’ own resources. As a result, users often create and configure digital artifacts in ways that are specialized for the unique infrastructure at a given HPC center. Each user of that center will face similar challenges as they develop specialized solutions to take full advantages of the center’s resources, potentially resulting in significant duplication of effort. Much duplicated effort could be avoided, however, if users of these centers found it easier to discover others’ solutions and artifacts as well as share their own. The FAIR principles address this problem by presenting guidelines focused around metadata practices to be implemented by vaguely defined “communities”; in practice, these tend to gather by domain (e.g. bioinformatics, geosciences, agriculture). Domain-based communities can unfortunately end up functioning as silos that tend both to inhibit sharing of solutions and best practices as well as to encourage fragile and unsustainable improvised solutions in the absence of best-practice guidance. We propose that these communities pursuing “science at scale” be nurtured both individually and collectively by HPC centers so that users can take advantage of shared challenges across disciplines and potentially across HPC centers. We describe an architecture based on the EOSC-Life FAIR Workflows Collaboratory, specialized for use with and inside HPC centers such as the Oak Ridge Leadership Computing Facility (OLCF), and we speculate on user incentives to encourage adoption. We note that a focus on FAIR workflow components rather than FAIR workflows is more likely to benefit the users of HPC centers.

Wilkinson, Sean [ORNL] (ORCID:0000000214437479)↗

BLDAP Intro to Python/Data Science Curriculum v1

The Github repository contains the Jupyter notebooks for the intro to Python / Data Science course for Berkeley Lab Director's Apprenticeship Program (BLDAP). This course is designed for students with little to no experience in coding to learn skills in Python necessary for data science. Students utilize Jupyter notebooks throughout the course. The overall goal is for students to learn how to use Python to clean, analyze, and visualize large data sets in order to communicate effectively their conclusions about the data set. Students apply the skills they learned on actual data sets provided by researchers in Berkeley Lab.

Hales, Laurel [Lawrence Berkeley National Laborato↗

Materials data science using CRADLE: A distributed, data-centric approach

Abstract There is a paradigm shift towards data-centric AI, where model efficacy relies on quality, unified data. The common research analytics and data lifecycle environment (CRADLE™) is an infrastructure and framework that supports a data-centric paradigm and materials data science at scale through heterogeneous data management, elastic scaling, and accessible interfaces. We demonstrate CRADLE’s capabilities through five materials science studies: phase identification in X-ray diffraction, defect segmentation in X-ray computed tomography, polymer crystallization analysis in atomic force microscopy, feature extraction from additive manufacturing, and geospatial data fusion. CRADLE catalyzes scalable, reproducible insights to transform how data is captured, stored, and analyzed. Graphical abstract

97 MATHEMATICS AND COMPUTING↗

A Tutorial Set to Prepare for Science with the Vera C. Rubin Observatory

In this poster the Rubin Observatory's Community Science team (CST) presents its current suite of tutorials, which are designed to help people make use of simulated data sets in preparation for the upcoming Legacy Survey of Space and Time (LSST). We will show examples of the tutorial contents, provide custom learning modules for different astronomical fields, and describe the online environment for data analysis (the Rubin Science Platform; RSP). We will also supply a checklist for how to obtain an RSP account and access the tutorials. All are welcome to drop by the poster or the Rubin booth in the exhibit hall with questions.

79 ASTRONOMY AND ASTROPHYSICS↗

Science & Technology Review: October-November 2024 DarkStar

At Lawrence Livermore National Laboratory, we focus on science and technology research to ensure our nation’s security. We also apply that expertise to solve other important national problems in energy, bioscience, and the environment. Science & Technology Review is published eight times a year to communicate, to a broad audience, the Laboratory’s scientific and technological accomplishments in fulfilling its primary missions. The publication’s goal is to help readers understand these accomplishments and appreciate their value to the individual citizen, the nation, and the world.

99 GENERAL AND MISCELLANEOUS↗

Measurements and Analyses to Enable Science for the Eastern Pacific Cloud Aerosol Precipitation Experiment (EPCAPE)

Coastal cities provide the opportunity to characterize the substantial effects of manmade particles on marine cloud properties and processes. La Jolla lies to the north of San Diego, California, but it is often about a day directly downwind of the major pollution sources located in the ports of Los Angeles and Long Beach. The large dynamic range of aerosol particle concentrations combined with the multi-hour to multi-day persistence of stratocumulus cloud layers makes the site ideal for investigating the seasonal changes in cloud and aerosol properties as well as the quantitative relationships between cloud and aerosol properties. The Eastern Pacific Cloud Aerosol Precipitation Experiment (EPCAPE) characterized the extent, radiative properties, aerosol interactions, and precipitation characteristics of stratocumulus clouds in the Eastern Pacific across all four seasons at two coastal sites in La Jolla. This project was designed to enhance and expand the scientific uses of the ARM AMF1 measurements from EPCAPE. The goal was to ensure meaningful observations were collected that would enable science. The project included the following objectives: (1) Reviewing the ARM AMF1 measurements [ARM, 2021a; b] and circulating a summary of the measurements each week, (2) Comparing AMF1 measurements to those provided by collaborators (including filter measurements at the pier), (3) Collecting and analyzing filter samples from Scripps Pier by Fourier Transform Infrared spectroscopy (FTIR) and X-ray Fluorescence (XRF), (4) Assisting in operations of instruments provided by Guest PIs when possible, and (5) Providing an initial compilation of EPCAPE aerosol and cloud seasonal differences. The expected outcomes of these objectives were enhanced proposals and publications using EPCAPE measurements by helping to identify instrument issues, expanded data access and awareness by distributing weekly plots and related summaries, improved source-related attribution of aerosols with elemental tracers, additional observations provided by Guest PIs, and accelerated ACI studies enabled by the compiled seasonal summaries of aerosol and cloud properties. Three examples of the science enabled by this project are findings that (i) aerosol and cloud aqueous production contributes more than half of sulfate particle mass concentration, (ii) upwind sources make chemical composition very similar at nearby sites despite local differences in meteorology, and (iii) most of the large mass concentration of semi-volatile organic components is co-emitted and co-evaporated with nitrate. Together these findings illustrate how ARM extended field campaigns in coastal regions can be used to constrain ACI processes with direct observations. By using the unique ARM suite of cloud radiative products in addition to the measured aerosol properties at a coastal location, we were able to address the more specific question of which aerosol particles cause how much of the effects on clouds. Identifying this signature in coastal areas provides an opportunity to test the representation of aerosol sources by global models in a range of clean and urban-influenced conditions.

Russell, Lynn [Univ. of California, San Diego, CA ↗