Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data mining”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Learning Global Proliferation Expertise Evolution Using AI-Driven Analytics and Public Information

Detecting and anticipating global proliferation expertise and capability evolution from unstructured, noisy, and incomplete public data streams is a highly desired, but extremely challenging task. Here, in this article, we present our pioneering data-driven approach to support the non-proliferation mission to detect and explain the evolution of proliferation expertise and capability development globally from terabytes of publicly available information (PAI), focusing on our knowledge extraction pipeline and descriptive analytics. We first discuss how we fuse nine open-source data streams, including multilingual data, to convert 4 TB of unstructured data to structured knowledge and encode dynamically evolving proliferation expertise representations—content and context graphs. For this, we rely on natural language processing (NLP) and deep learning (DL) models to perform information extraction, topic modeling, and distributed text representation (aka embedding) learning. We then present interactive, usable, and explainable descriptive analytics to refine domain knowledge and present it in a human-understandable form. Finally, we introduce future work avenues that will leverage our dynamic knowledge representations and descriptive analytics to enable predictive and prescriptive inferences to achieve real-time domain understanding and contextual reasoning about global proliferation expertise and capability evolution.

98 NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL P↗

Wide-ranging predictions of new stable compounds powered by recommendation engines

The computational search for new stable inorganic compounds is faster than ever, thanks to high-throughput density functional theory (DFT). However, stable compound searches remain highly expensive because of the enormous search space and the cost of DFT calculations. To aid these searches, recommendation engines have been developed. We conduct a systematic comparison of the performance of previously developed recommendation engines, specifically ones based on elemental substitution, data mining, and neural network prediction of formation enthalpy. After identifying ways to improve the recommendation engines, we find the neural network to be superior at recommending stable Heusler compounds. Armed with improved recommendation engines, we identify tens of thousands of compounds that are stable at zero temperature and pressure, now available in the Open Quantum Materials Database. We summarize this diverse pool of compounds, including the elusive mixed anion compounds, and two of their many applications: thermoelectricity and solar thermochemical fuel production.

Science & Technology - Other Topics↗

Fragile Earth: AI for Climate Mitigation, Adaptation, and Environmental Justice

The Fragile EarthWorkshop is a recurring event that gathers the research community to find and explore howdata science can measure and progress climate and social issues, following the framework of the United Nations Sustainable Development Goals (SDGs).Fragile Earth 2022: AI for Climate Mitigation, Adaptation, and Environmental Justice is a workshop taking place as part of the ACM's KDD 2022 Conference on research in knowledge discovery and data mining and their applications. The dates for the Conference are August 14-18, 2022.

Abe, Naoki↗

Noise-Resilient and Reduced Depth Approximate Adders for NISQ Quantum Computing

The "Noisy intermediate-scale quantum" NISQ machine era primarily focuses on mitigating noise, controlling errors, and executing high-fidelity operations, hence requiring shallow circuit depth and noise robustness. Approximate computing is a novel computing paradigm that produces imprecise results by relaxing the need for fully precise output for error-tolerant applications including multimedia, data mining, and image processing. We investigate how approximate computing can improve the noise resilience of quantum adder circuits in NISQ quantum computing. We propose five designs of approximate quantum adders to reduce depth while making them noise-resilient, in which three designs are with carryout, while two are without carryout. We have used novel design approaches that include approximating the Sum only from the inputs (pass-through designs) and having zero depth, as they need no quantum gates. The second design style uses a single CNOT gate to approximate the SUM with a constant depth of O(1). We performed our experimentation on IBM Qiskit on noise models including thermal, depolarizing, amplitude damping, phase damping, and bitflip: (i) Compared to exact quantum ripple carry adder without carryout the proposed approximate adders without carryout have improved fidelity ranging from 8.34% to 219.22%, and (ii) Compared to exact quantum ripple carry adder with carryout the proposed approximate adders with carryout have improved fidelity ranging from 8.23% to 371%. Further, the proposed approximate quantum adders are evaluated in terms of various error metrics.

Gaur, Bhaskar↗

Studying CPU and memory utilization of applications on Fujitsu A64FX and Nvidia Grace Superchip

ARM-based manycore CPU architectures are well-positioned to provide the rising memory throughput requirements of modern data intensive scientific applications in High Performance Computing (HPC). The Fujitsu A64FX CPU platform is based on the ARM v8.2A architecture, and is the processor of the flagship Japanese supercomputer - "Fugaku", which was previously ranked as the #1 supercomputer in the world according to the Top500 list. The Nvidia Grace superchip features 144 Neoverse V2 cores based on the ARMv9 architecture with 4x128b SVE2, providing exceptional computational power. The chip supports up to 480GB of memory, making it ideal for AI, machine learning, and scientific computing workloads. In this paper, we conduct a thorough performance exploration of a variety of parallel bandwidth-sensitive benchmarks and applications compiled with the native Fujitsu compiler on a Fugaku A64FX compute node and ARM (LLVM) Compiler on an NVIDIA Grace superchip compute node, engaging all the computational cores per cluster using OpenMP multithreading (assuming the cores can drive the available bandwidth). Our ultimate goals are to study the resource utilization of scientific applications and benchmarks on A64FX and Grace superchip, considering graph application scenarios ( GAP Benchmark suite) and eleven appli- cation proxies from the Rodinia heterogeneous benchmark suite (considering domains such as Data Mining, Bioinformatics, Fluid Dynamics, Pattern Recognition, etc.). Through exhaustive performance monitoring, we quantify the resource utilization of diverse OpenMP-based HPC applications on both the Fujitsu A64FX and the Nvidia Grace Superchip platforms.

benchmarking, Performance Analysis, High performan↗

TEAL

TEAL is a financial performance calculator plugin for the RAVEN code, framework, resolving around the computation of Net Present Value and associated financial metrics. TEAL can make use of inflation rates, taxation, escalation factors, capital expenditure economy of scale scaling factors. The unique feature of TEAL is the capability to be linked with RAVEN external models and build corresponding cash flows using the variables computed by those external models. In addition to be able to use the capability to generate cash flows derived from complex physical models generated by RAVEN, another distinctive feature of TEAL is the capability to provide financial risk/probabilistic metrics that can empower RAVEN to perform optimization/analysis driven by financial risk augmentations. Optimization, robust optimization, parametric studies, large parallel simulations, sensitivity analysis, data mining, etc. are just some of the capabilities that can be leveraged.

Alfonsi, Andrea↗

Reposcanner

SAND2023-05455O Reposcanner provides a highly modular, extensible framework for defining routines for mining data from software repositories and performing analyses on that data to yield valuable insights on team behaviors. Reposcanner features seamless support for different version control platforms like GitHub, Gitlab, and Bitbucket; smart parsing of URLs; intelligent credential management capabilities; and a comprehensive test suite. Reposcanner is connected to the Exascale Computing Project and is intended for research purposes.

Mundt, Miranda↗

Perspectives on AI Architectures and Codesign for Earth System Predictability

Abstract Recently, the U.S. Department of Energy (DOE), Office of Science, Biological and Environmental Research (BER), and Advanced Scientific Computing Research (ASCR) programs organized and held the Artificial Intelligence for Earth System Predictability (AI4ESP) workshop series. From this workshop, a critical conclusion that the DOE BER and ASCR community came to is the requirement to develop a new paradigm for Earth system predictability focused on enabling artificial intelligence (AI) across the field, laboratory, modeling, and analysis activities, called model experimentation (ModEx). BER’s ModEx is an iterative approach that enables process models to generate hypotheses. The developed hypotheses inform field and laboratory efforts to collect measurement and observation data, which are subsequently used to parameterize, drive, and test model (e.g., process based) predictions. A total of 17 technical sessions were held in this AI4ESP workshop series. This paper discusses the topic of the AI Architectures and Codesign session and associated outcomes. The AI Architectures and Codesign session included two invited talks, two plenary discussion panels, and three breakout rooms that covered specific topics, including 1) DOE high-performance computing (HPC) systems, 2) cloud HPC systems, and 3) edge computing and Internet of Things (IoT). We also provide forward-looking ideas and perspectives on potential research in this codesign area that can be achieved by synergies with the other 16 session topics. These ideas include topics such as 1) reimagining codesign, 2) data acquisition to distribution, 3) heterogeneous HPC solutions for integration of AI/ML and other data analytics like uncertainty quantification with Earth system modeling and simulation, and 4) AI-enabled sensor integration into Earth system measurements and observations. Such perspectives are a distinguishing aspect of this paper. Significance Statement This study aims to provide perspectives on AI architectures and codesign approaches for Earth system predictability. Such visionary perspectives are essential because AI-enabled model-data integration has shown promise in improving predictions associated with climate change, perturbations, and extreme events. Our forward-looking ideas guide what is next in codesign to enhance Earth system models, observations, and theory using state-of-the-art and futuristic computational infrastructure.

54 ENVIRONMENTAL SCIENCES↗

Quantitative gas-phase transmission electron microscopy: Where are we now and what comes next?

Abstract Based on historical developments and the current state of the art in gas-phase transmission electron microscopy (GP-TEM), we provide a perspective covering exciting new technologies and methodologies of relevance for chemical and surface sciences. Considering thermal and photochemical reaction environments, we emphasize the benefit of implementing gas cells, quantitative TEM approaches using sensitive detection for structured electron illumination (in space and time) and data denoising, optical excitation, and data mining using autonomous machine learning techniques. These emerging advances open new ways to accelerate discoveries in chemical and surface sciences. Graphical abstract

36 MATERIALS SCIENCE↗

Nuclear Physics Exascale Requirements Review: An Office of Science Review sponsored jointly by Advanced Scientific Computing Research and Nuclear Physics, June 15 - 17, 2016, Gaithersburg, Maryland

Imagine being able to predict — with unprecedented accuracy and precision — the structure of the proton and neutron, and the forces between them, directly from the dynamics of quarks and gluons, and then using this information in calculations of the structure and reactions of atomic nuclei and of the properties of dense neutron stars (NSs). Also imagine discovering new and exotic states of matter, and new laws of nature, by being able to collect more experimental data than we dream possible today, analyzing it in real time to feed back into an experiment, and curating the data with full tracking capabilities and with fully distributed data mining capabilities. Making this vision a reality would improve basic scientific understanding, enabling us to precisely calculate, for example, the spectrum of gravity waves emitted during NS coalescence, and would have important societal applications in nuclear energy research, stockpile stewardship, and other areas. This review presents the components and characteristics of the exascale computing ecosystems necessary to realize this vision.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Exploiting Commonly-Reported Age-Hardening Data and Discovering Systematics Across Metallic Alloys and Alloy Systems to Identify Corrosion’s Most Influential Factors

This work explored ways to data mine legacy literature and predict solid-solid precipitation in support of simpler assessment of corrosion propensity in metallic alloys. Of interest was locating the peak age watershed (maximum hardness or strength), beyond which lies the regime termed “overaging.” A diligent search of literature and reference books, discussions with SMEs, and application of various algorithms showed that none of the premises going in held up to scrutiny.

36 MATERIALS SCIENCE↗

Science Breakthroughs 2030. Final report

Agriculture is a fundamental societal activity, characterized by many different landscapes, crops, markets, and participants. Food, agricultural, and biofuels products are central to the daily life of all citizens, though most do not recognize the fragility of the environment that brings forth this abundance. As is noted in the 2012 report from the President's Council of Advisors on Science and Technology, Agricultural Preparedness and the United States Agricultural Research Enterprise (PCAST, 2012) the food and agricultural system faces constant challenges in: Managing new pests, pathogens, and invasive plants. Increasing the efficiency of water use. Growing food in a changing climate. Reducing the environmental footprint of agriculture. Managing the production of bioenergy. Producing safe and nutritious food. Assisting with global food security and maintaining abundant yields. Science Breakthroughs 2030 was organized to identify the most compelling research directions in food and agriculture, in particular those empowered by the application of insights and tools from disciplines of science and engineering not typically associated with food and agricultural research. A committee appointed by the Chairman of the National Research Council explored ideas for research directions with input from the scientific community, with the objective of producing a report describing ambitious and achievable scientific pathways to address major problems and create new opportunities in food and agriculture. Following numerous meetings, a jamboree, and town hall, the appointed committee prepared a report that has subsequently become a reference for federal agencies supporting research in the food and agricultural space. It highlights five key areas for research investment with broad application across food and agriculture: integrated systems research; sensor development; data mining and information sciences, genomics; and the microbiome.

09 BIOMASS FUELS↗

Status Report on the Molten Salt Thermodynamic Database (MSTDB) Development (FY20)

The current status of the Molten Salt Thermodynamic Database (MSTDB) is reported. While a series of informational exchange meetings with Moten Salt Reactor (MSR) developers has given the NEAMS program insight into the systems of interest, the current effort is devoted to mining data already available in the literature for continued development of MSTDB. Many relevant systems have not been previously studied. Therefore, once the data from the literature is completely reviewed, curated, and used as inputs for MSTDB, new thermodynamic values will need to be generated using computational and experimental approaches.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

LABS: Los Alamos Benchmark Suite - Current status and planning [Slides]

The LABS repository now has over 500 verified and revised input files with about 600 input files left to get up to speed with what we currently use. After that there are many, many more to go. We're in the process of building code infrastructure, moving to template output files, automating input file generation, outputting file data mining for calculation results, and benchmarking selection on sensitivity and other benchmark characteristics. We intend to open source LABS. We'll most likely release the input files sooner rather than later. The associated tool will probably be open sourced later.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Using Machine Learning to Identify Optimum Prospects for Offshore Geologic CO 2 Storage and Enhanced Oil Recovery in the Gulf of Mexico

The SECARB Offshore Partnership is a government-industry partnership focused on assembling the knowledge base required for secure, long-term, large-scale CO 2 subsea storage in saline formations and mature oil reservoirs. The SECARB Offshore Partnership is evaluating potential storage opportunities in the Central and Eastern planning areas of the outer continental shelf (OCS) of the Gulf of Mexico. The evaluation focuses on active and depleted oil and gas fields and potentially associated CO 2 -enhanced oil recovery (CO 2 -EOR), as well as, deep saline storage resources. As part of Subtask 3.0 (Offshore Storage Resources Characterization) and Subtask 4.0 (Risk Assessment, Simulation, and Modeling), Oklahoma State University (OSU) is developing a machine learning system that employs the SAS Institute’s Viya platform with visual data mining and machine learning to identify and characterize optimum prospects for enhanced oil recovery and stacked storage in saline reservoirs. This report summarizes this effort to date and discusses the many variables that can be used to characterize these storage objectives, as well as, the design vision of the Viya machine learning system.

02 PETROLEUM↗

COVID-19 Data Curation Effort: An Initial Analysis of the Data

During the COVID-19 pandemic of 2020, major case reporting outlets quickly coalesced around two or three primary vendors. Johns Hopkins University and The New York Times were among the more prominent, and all were of great value to the nation, particularly during the uncertain early stages of the pandemic. They primarily focused on three major attributes: number of new cases, deaths, and recovery. Recognizing that many states were reporting very detailed data sets (e.g., hospital beds) at a county level or finer, the ORNL Pandemic Modeling team embarked on a major data curation effort from March to June 2020 for the purpose of capturing this wealth of detailed data. The challenge of curating this data was daunting. The number of attributes reported by the states grew on almost on a weekly basis. States were routinely shifting their web tool strategies away from easily parsable HTML-based formatting to new Tableau and ArcGIS content. This growth in the sheer number of attributes, combined with the unpredictable shifts in data format, meant an aggressive and agile combination of automated scripting and manual scraping was required to capture new daily streams. Further, the team had to scale up staff and widen its approach for capture and storage. As a result, the team collected more than 11 million data points. Following the close of this data collection effort on June 30 th , 2020, the team embarked on a major effort to appraise what had been collected, including an inventory list, spatial completeness, temporal completeness, scale and geographic characteristics, and a determination. A report on this matter was submitted on September 15 th , 2020, titled “DOE COVID-19 Data Curation Effort: Overview of Data Collection Coverage”. Over 2000 unique attributes had been netted over a wide range of spatial scales, including state, county, zip codes, health regions, and census blocks. Over 11 million individual data points were collected across these attributes, and spatial coverage (in total) included all 50 states and multiple territories. What became apparent in the process is that in the absence of any data standards, many states reported a wide variety of unique attributes that were not always compatible with attributes reported in other states. As time continued, states began adding new attributes and offering finer grain detail in some older attributes. This meant that not all data streams existed for the entire time period; in fact, the number tended to increase dramatically towards the end. Often, states would begin an attribute series and then stop altogether. These highly variable and uncertain conditions illuminated the need for harmonization approaches that would reconcile and conflate changing attribute names and detail over time. For example, grouping racial data reported as either Black or African American, depending on the state, into a single harmonized attribute. These choices would make a within-state analysis possible during the time period and lead to potential between-state analytics later on. This was almost entirely a manual decision process, requiring some subjective decision-making at times, to prevent a fragmented, short-lived collection of time series fragments that would offer few insights into trends, patterns, and correlates. This report imports harmonized data for state and county into the World Spatio-Temporal Analytics and Mapping Project (WSTAMP). WSTAMP is a major space-time analysis and visualization tool developed at ORNL for the National Geospatial-Intelligence Agency specifically for this kind of exploratory analysis. WSTAMP offers a rich analytical and graphical environment consisting of a wide range of analytics. These include time series plots, statistical summaries, data mining techniques, trend and pattern detection, and hypothesis generation.

59 BASIC BIOLOGICAL SCIENCES↗

HL-LHC Analysis With ROOT: ROOT Project Input to the HL-LHC Computing Review (Stage 2)

ROOT is high energy physics' software for storing and mining data in a statistically sound way, to publish results with scientific graphics. It is evolving since 25 years, now providing the storage format for more than one exabyte of data; virtually all high energy physics experiments use ROOT. With another significant increase in the amount of data to be handled scheduled to arrive in 2027, ROOT is preparing for a massive upgrade of its core ingredients. As part of a review of crucial software for high energy physics, the ROOT team has documented its R&D plans for the coming years.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Preliminary Report on Compositional Specification for Printed 316SS

The Advanced Materials and Manufacturing Technologies Program is focused on accelerating the development and deployment of advanced materials and components fabricated via additive manufacturing (AM). A key challenge for the widespread adoption of AM is the variability in material properties, which can originate from powder feedstock, process parameters, geometry, and the machine. AM powder feedstock compositions are based on those developed for conventional manufacturing. However, the smaller melt pools combined with multiple melt cycles over the course of a build can have substantial effects on the local phase selection resulting from minor batch-to-batch feedstock variability. Therefore, this report focuses on coupling CALPHAD calculations with data mining and visualization to aid the development of compositional specifications for stainless steel 316, an important alloy for nuclear applications and suitable for AM. This work reports that δ-ferrite formation has a stronger dependence on the amount of nickel in the alloy despite nickel being an austenite stabilizer. These findings show the need for exercising a tighter control on the amount of nickel in the alloy compared with the amount of chromium, which is a ferrite stabilizer.

36 MATERIALS SCIENCE↗