Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data mining”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17

Simulating single-particle dynamics in magnetized plasmas: The RMF code

The RMF (Rotating Magnetic Field) code is designed to calculate the motion of a charged particle in a given electromagnetic field. It integrates Hamilton’s equations in cylindrical coordinates using an adaptive predictor-corrector double-precision variable-coefficient ordinary differential equation solver for speed and accuracy. RMF has multiple capabilities for the field. Particle motion is initialized by specifying the position and velocity vectors. Here, the six-dimensional state vector and derived quantities are saved as functions of time. A post-processing graphics code, XDRAW, is used on the stored output to plot up to 12 windows of any two quantities using different colors to denote successive time intervals. Multiple cases of RMF may be run in parallel and perform data mining on the results. Recent features are a synthetic diagnostic for simulating the observations of charge-exchange-neutral energy distributions and RF grids to explore a Fermi acceleration parallel to static magnetic fields.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

"BAAE" instabilities observed without fast ion drive

The instability that was previously identified as a fast-ion driven beta-induced Alfv´en-acoustic eigenmode (BAAE) in DIII-D was misidentified. In a dedicated experiment, low frequency modes (LFM) with characteristic “Christmas light” patterns of brief instability linked to the safety factor evolution occur in plasmas with electron temperature T e ≳ 2.1 keV but modest beta. To isolate the importance of different driving gradients on these modes, the electron cyclotron heating power and 80 keV, sub-Alfv´enic neutral beams are altered for 50-100 ms durations in reproducible discharges. Although beta-induced Alfv´en eigenmodes and reversed-shear Alfv´en eigenmodes stabilize when beam injection ceases (as expected for a fast-ion driven instability), the low frequency modes that were called BAAEs persist. Data mining reveals that characteristic LFM instabilities can occur in discharges with no beam heating but strong electron cyclotron heating. A large database of over 1000 discharges shows that LFMs are only unstable in plasmas with hot electrons but modest overall beta. The experimental LFMs have low frequencies (comparable to diamagnetic drift frequencies) in the plasma frame, occur near the minimum of the safety factor q min , and appear when q min is close to rational values. In conclusion, theoretical analysis suggests that the LFMs are a low frequency reactive instability of predominately Alfv´enic polarization.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Electron-phonon coupling strength from ab initio frozen-phonon approach

In this work, we propose a fast method for high-throughput screening of potential superconducting materials. The method is based on calculating metallic screening of zone-center phonon modes, which provides an accurate estimate for the electron-phonon coupling strength. This method is complementary to the recently proposed Rigid Muffin Tin (RMT) method, which amounts to integrating the electronphonon coupling over the entire Brillouin zone (as opposed to the zone center), but in a relatively inferior approximation. We illustrate the use of this method by applying it to MgB 2 , where the high-temperature superconductivity is known to be driven largely by the zone-center modes, and compare it to a sister compound AlB 2 . We further illustrate the usage of this descriptor by screening a large number of binary hydrides, for which accurate first-principle calculations of electron-phonon coupling have been recently published. Together with the RMT descriptor, this method opens a way to perform initial high-throughput screening in search of conventional superconductors via machine learning or data mining.

36 MATERIALS SCIENCE↗

SymProp: Scaling Sparse Symmetric Tucker Decomposition via Symmetry Propagation

Sparse symmetric tensors are an important class of tensors, and their decompositions serve as powerful tools for revealing low-rank structures. This paper introduces SymProp, a novel approach for scaling sparse symmetric Tucker decomposition by propagating symmetry through intermediate computations. SymProp optimizes two key computational kernels: Sparse Symmetric Tensor Times Same Matrix chain (S3 TTMc) for Higher-Order Orthogonal Iteration (HOOI) and Sparse Symmetric Tensor Times Same Matrix chain Times Core (S3 TTMcTC) for Higher-Order QR Iteration (HOQRI). Our method employs a metaprogramming-based index iteration approach to efficiently handle the upper triangular parts of intermediate dense symmetric tensors. SymProp achieves up to 50.9× speedup over SPLATT and up to 360.8× over Compressed Sparse Symmetric (CSS) format on the S3 TTMc operation. Moreover, our S3 TTMc and S3 TTMcTC implementations support tensor orders four levels higher than state-of-the-art methods. Our HOQRI demonstrates superior scalability and up to a 33.6× speedup over optimized HOOI. By enabling more scalable Tucker decompositions for higher orders, decomposition ranks, and dimension sizes, SymProp opens new possibilities for analyzing complex hypergraph structures in fields such as network science, data mining, and machine learning.

Li, Zecheng [North Carolina State University]↗

Scalable Knowledge Graph Analytics at 136 Petaflop/s

We are motivated by newly proposed methods for data mining large-scale corpora of scholarly publications, such as the full biomedical literature, which may consist of tens of millions of papers spanning decades of research. In this setting, analysts seek to discover how concepts relate to one another. They construct graph representations from annotated text databases and then formulate the relationship-mining problem as one of computing all-pairs shortest paths (APSP), which becomes a significant bottleneck. In this context, we present a new high-performance algorithm and implementation of the Floyd-Warshall algorithm for distributed-memory parallel computers accelerated by GPUs, which we call DSNAPSHOT (Distributed Accelerated Semiring All-Pairs Shortest Path). For our largest experiments, we ran DSNAPSHOT on a connected input graph with millions of vertices using 4, 096nodes (24,576GPUs) of the Oak Ridge National Laboratory's Summit supercomputer system. We find DSNAPSHOT achieves a sustained performance of 136×1015 floating-point operations per second (136petaflop/s) at a parallel efficiency of 90% under weak scaling and, in absolute speed, 70% of the best possible performance given our computation (in the single-precision tropical semiring or “min-plus” algebra). Looking forward, we believe this novel capability will enable the mining of scholarly knowledge corpora when embedded and integrated into artificial intelligence-driven natural language processing workflows at scale.

Kannan, Ramakrishnan {ramki}↗

Learning Global Proliferation Expertise Evolution Using AI-Driven Analytics and Public Information

Detecting and anticipating global proliferation expertise and capability evolution from unstructured, noisy, and incomplete public data streams is a highly desired, but extremely challenging task. Here, in this article, we present our pioneering data-driven approach to support the non-proliferation mission to detect and explain the evolution of proliferation expertise and capability development globally from terabytes of publicly available information (PAI), focusing on our knowledge extraction pipeline and descriptive analytics. We first discuss how we fuse nine open-source data streams, including multilingual data, to convert 4 TB of unstructured data to structured knowledge and encode dynamically evolving proliferation expertise representations—content and context graphs. For this, we rely on natural language processing (NLP) and deep learning (DL) models to perform information extraction, topic modeling, and distributed text representation (aka embedding) learning. We then present interactive, usable, and explainable descriptive analytics to refine domain knowledge and present it in a human-understandable form. Finally, we introduce future work avenues that will leverage our dynamic knowledge representations and descriptive analytics to enable predictive and prescriptive inferences to achieve real-time domain understanding and contextual reasoning about global proliferation expertise and capability evolution.

98 NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL P↗

Wide-ranging predictions of new stable compounds powered by recommendation engines

The computational search for new stable inorganic compounds is faster than ever, thanks to high-throughput density functional theory (DFT). However, stable compound searches remain highly expensive because of the enormous search space and the cost of DFT calculations. To aid these searches, recommendation engines have been developed. We conduct a systematic comparison of the performance of previously developed recommendation engines, specifically ones based on elemental substitution, data mining, and neural network prediction of formation enthalpy. After identifying ways to improve the recommendation engines, we find the neural network to be superior at recommending stable Heusler compounds. Armed with improved recommendation engines, we identify tens of thousands of compounds that are stable at zero temperature and pressure, now available in the Open Quantum Materials Database. We summarize this diverse pool of compounds, including the elusive mixed anion compounds, and two of their many applications: thermoelectricity and solar thermochemical fuel production.

Science & Technology - Other Topics↗

Fragile Earth: AI for Climate Mitigation, Adaptation, and Environmental Justice

The Fragile EarthWorkshop is a recurring event that gathers the research community to find and explore howdata science can measure and progress climate and social issues, following the framework of the United Nations Sustainable Development Goals (SDGs).Fragile Earth 2022: AI for Climate Mitigation, Adaptation, and Environmental Justice is a workshop taking place as part of the ACM's KDD 2022 Conference on research in knowledge discovery and data mining and their applications. The dates for the Conference are August 14-18, 2022.

Abe, Naoki↗

Noise-Resilient and Reduced Depth Approximate Adders for NISQ Quantum Computing

The "Noisy intermediate-scale quantum" NISQ machine era primarily focuses on mitigating noise, controlling errors, and executing high-fidelity operations, hence requiring shallow circuit depth and noise robustness. Approximate computing is a novel computing paradigm that produces imprecise results by relaxing the need for fully precise output for error-tolerant applications including multimedia, data mining, and image processing. We investigate how approximate computing can improve the noise resilience of quantum adder circuits in NISQ quantum computing. We propose five designs of approximate quantum adders to reduce depth while making them noise-resilient, in which three designs are with carryout, while two are without carryout. We have used novel design approaches that include approximating the Sum only from the inputs (pass-through designs) and having zero depth, as they need no quantum gates. The second design style uses a single CNOT gate to approximate the SUM with a constant depth of O(1). We performed our experimentation on IBM Qiskit on noise models including thermal, depolarizing, amplitude damping, phase damping, and bitflip: (i) Compared to exact quantum ripple carry adder without carryout the proposed approximate adders without carryout have improved fidelity ranging from 8.34% to 219.22%, and (ii) Compared to exact quantum ripple carry adder with carryout the proposed approximate adders with carryout have improved fidelity ranging from 8.23% to 371%. Further, the proposed approximate quantum adders are evaluated in terms of various error metrics.

Gaur, Bhaskar↗

Studying CPU and memory utilization of applications on Fujitsu A64FX and Nvidia Grace Superchip

ARM-based manycore CPU architectures are well-positioned to provide the rising memory throughput requirements of modern data intensive scientific applications in High Performance Computing (HPC). The Fujitsu A64FX CPU platform is based on the ARM v8.2A architecture, and is the processor of the flagship Japanese supercomputer - "Fugaku", which was previously ranked as the #1 supercomputer in the world according to the Top500 list. The Nvidia Grace superchip features 144 Neoverse V2 cores based on the ARMv9 architecture with 4x128b SVE2, providing exceptional computational power. The chip supports up to 480GB of memory, making it ideal for AI, machine learning, and scientific computing workloads. In this paper, we conduct a thorough performance exploration of a variety of parallel bandwidth-sensitive benchmarks and applications compiled with the native Fujitsu compiler on a Fugaku A64FX compute node and ARM (LLVM) Compiler on an NVIDIA Grace superchip compute node, engaging all the computational cores per cluster using OpenMP multithreading (assuming the cores can drive the available bandwidth). Our ultimate goals are to study the resource utilization of scientific applications and benchmarks on A64FX and Grace superchip, considering graph application scenarios ( GAP Benchmark suite) and eleven appli- cation proxies from the Rodinia heterogeneous benchmark suite (considering domains such as Data Mining, Bioinformatics, Fluid Dynamics, Pattern Recognition, etc.). Through exhaustive performance monitoring, we quantify the resource utilization of diverse OpenMP-based HPC applications on both the Fujitsu A64FX and the Nvidia Grace Superchip platforms.

benchmarking, Performance Analysis, High performan↗

TEAL

TEAL is a financial performance calculator plugin for the RAVEN code, framework, resolving around the computation of Net Present Value and associated financial metrics. TEAL can make use of inflation rates, taxation, escalation factors, capital expenditure economy of scale scaling factors. The unique feature of TEAL is the capability to be linked with RAVEN external models and build corresponding cash flows using the variables computed by those external models. In addition to be able to use the capability to generate cash flows derived from complex physical models generated by RAVEN, another distinctive feature of TEAL is the capability to provide financial risk/probabilistic metrics that can empower RAVEN to perform optimization/analysis driven by financial risk augmentations. Optimization, robust optimization, parametric studies, large parallel simulations, sensitivity analysis, data mining, etc. are just some of the capabilities that can be leveraged.

Alfonsi, Andrea↗

Reposcanner

SAND2023-05455O Reposcanner provides a highly modular, extensible framework for defining routines for mining data from software repositories and performing analyses on that data to yield valuable insights on team behaviors. Reposcanner features seamless support for different version control platforms like GitHub, Gitlab, and Bitbucket; smart parsing of URLs; intelligent credential management capabilities; and a comprehensive test suite. Reposcanner is connected to the Exascale Computing Project and is intended for research purposes.

Mundt, Miranda↗

Perspectives on AI Architectures and Codesign for Earth System Predictability

Abstract Recently, the U.S. Department of Energy (DOE), Office of Science, Biological and Environmental Research (BER), and Advanced Scientific Computing Research (ASCR) programs organized and held the Artificial Intelligence for Earth System Predictability (AI4ESP) workshop series. From this workshop, a critical conclusion that the DOE BER and ASCR community came to is the requirement to develop a new paradigm for Earth system predictability focused on enabling artificial intelligence (AI) across the field, laboratory, modeling, and analysis activities, called model experimentation (ModEx). BER’s ModEx is an iterative approach that enables process models to generate hypotheses. The developed hypotheses inform field and laboratory efforts to collect measurement and observation data, which are subsequently used to parameterize, drive, and test model (e.g., process based) predictions. A total of 17 technical sessions were held in this AI4ESP workshop series. This paper discusses the topic of the AI Architectures and Codesign session and associated outcomes. The AI Architectures and Codesign session included two invited talks, two plenary discussion panels, and three breakout rooms that covered specific topics, including 1) DOE high-performance computing (HPC) systems, 2) cloud HPC systems, and 3) edge computing and Internet of Things (IoT). We also provide forward-looking ideas and perspectives on potential research in this codesign area that can be achieved by synergies with the other 16 session topics. These ideas include topics such as 1) reimagining codesign, 2) data acquisition to distribution, 3) heterogeneous HPC solutions for integration of AI/ML and other data analytics like uncertainty quantification with Earth system modeling and simulation, and 4) AI-enabled sensor integration into Earth system measurements and observations. Such perspectives are a distinguishing aspect of this paper. Significance Statement This study aims to provide perspectives on AI architectures and codesign approaches for Earth system predictability. Such visionary perspectives are essential because AI-enabled model-data integration has shown promise in improving predictions associated with climate change, perturbations, and extreme events. Our forward-looking ideas guide what is next in codesign to enhance Earth system models, observations, and theory using state-of-the-art and futuristic computational infrastructure.

54 ENVIRONMENTAL SCIENCES↗

Quantitative gas-phase transmission electron microscopy: Where are we now and what comes next?

Abstract Based on historical developments and the current state of the art in gas-phase transmission electron microscopy (GP-TEM), we provide a perspective covering exciting new technologies and methodologies of relevance for chemical and surface sciences. Considering thermal and photochemical reaction environments, we emphasize the benefit of implementing gas cells, quantitative TEM approaches using sensitive detection for structured electron illumination (in space and time) and data denoising, optical excitation, and data mining using autonomous machine learning techniques. These emerging advances open new ways to accelerate discoveries in chemical and surface sciences. Graphical abstract

36 MATERIALS SCIENCE↗

Nuclear Physics Exascale Requirements Review: An Office of Science Review sponsored jointly by Advanced Scientific Computing Research and Nuclear Physics, June 15 - 17, 2016, Gaithersburg, Maryland

Imagine being able to predict — with unprecedented accuracy and precision — the structure of the proton and neutron, and the forces between them, directly from the dynamics of quarks and gluons, and then using this information in calculations of the structure and reactions of atomic nuclei and of the properties of dense neutron stars (NSs). Also imagine discovering new and exotic states of matter, and new laws of nature, by being able to collect more experimental data than we dream possible today, analyzing it in real time to feed back into an experiment, and curating the data with full tracking capabilities and with fully distributed data mining capabilities. Making this vision a reality would improve basic scientific understanding, enabling us to precisely calculate, for example, the spectrum of gravity waves emitted during NS coalescence, and would have important societal applications in nuclear energy research, stockpile stewardship, and other areas. This review presents the components and characteristics of the exascale computing ecosystems necessary to realize this vision.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Exploiting Commonly-Reported Age-Hardening Data and Discovering Systematics Across Metallic Alloys and Alloy Systems to Identify Corrosion’s Most Influential Factors

This work explored ways to data mine legacy literature and predict solid-solid precipitation in support of simpler assessment of corrosion propensity in metallic alloys. Of interest was locating the peak age watershed (maximum hardness or strength), beyond which lies the regime termed “overaging.” A diligent search of literature and reference books, discussions with SMEs, and application of various algorithms showed that none of the premises going in held up to scrutiny.

36 MATERIALS SCIENCE↗

Science Breakthroughs 2030. Final report

Agriculture is a fundamental societal activity, characterized by many different landscapes, crops, markets, and participants. Food, agricultural, and biofuels products are central to the daily life of all citizens, though most do not recognize the fragility of the environment that brings forth this abundance. As is noted in the 2012 report from the President's Council of Advisors on Science and Technology, Agricultural Preparedness and the United States Agricultural Research Enterprise (PCAST, 2012) the food and agricultural system faces constant challenges in: Managing new pests, pathogens, and invasive plants. Increasing the efficiency of water use. Growing food in a changing climate. Reducing the environmental footprint of agriculture. Managing the production of bioenergy. Producing safe and nutritious food. Assisting with global food security and maintaining abundant yields. Science Breakthroughs 2030 was organized to identify the most compelling research directions in food and agriculture, in particular those empowered by the application of insights and tools from disciplines of science and engineering not typically associated with food and agricultural research. A committee appointed by the Chairman of the National Research Council explored ideas for research directions with input from the scientific community, with the objective of producing a report describing ambitious and achievable scientific pathways to address major problems and create new opportunities in food and agriculture. Following numerous meetings, a jamboree, and town hall, the appointed committee prepared a report that has subsequently become a reference for federal agencies supporting research in the food and agricultural space. It highlights five key areas for research investment with broad application across food and agriculture: integrated systems research; sensor development; data mining and information sciences, genomics; and the microbiome.

09 BIOMASS FUELS↗

Status Report on the Molten Salt Thermodynamic Database (MSTDB) Development (FY20)

The current status of the Molten Salt Thermodynamic Database (MSTDB) is reported. While a series of informational exchange meetings with Moten Salt Reactor (MSR) developers has given the NEAMS program insight into the systems of interest, the current effort is devoted to mining data already available in the literature for continued development of MSTDB. Many relevant systems have not been previously studied. Therefore, once the data from the literature is completely reviewed, curated, and used as inputs for MSTDB, new thermodynamic values will need to be generated using computational and experimental approaches.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗