Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data mining”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

A High-Throughput Computing Infrastructure to Generate Custom, Open Community Geothermal Datasets

The most significant challenge facing geothermal research, development, and deployment is a lack of comprehensive datasets describing the geological and economical properties of North America. Automated knowledge base construction, the process of designing algorithms to analyze text and images to programmatically build new datasets, is one possible solution to this problem. The xDD library of full-text scientific articles (https://xdd.wisc.edu) is one of the largest collections of open and controlled-access scientific documents available for knowledge base construction in the world, but it has been underutilized by experts in geothermal research. The xDD development team attributed the lack of engagement by software developers and geothermal researchers to two perceived shortcomings of the system. First, the workflow for obtaining data from xDD for local development and testing of data mining applications was unnecessarily abstruse and required significant manual intervention by xDD systems administrators. Second, although xDD already held articles from a broad cross-section of scientific literature with an emphasis on the geosciences, it did not have an explicit set of geothermal research documents that could serve as the nucleus of a geothermal data mining application. To address these issues, the Automated Data Extraction PlaTform (ADEPT) was proposed to extend the data distribution capabilities of the xDD system. The ADEPT extension added the following four key features to xDD: 1) integration of National Geothermal Data System (NGDS) documents into the xDD library to provide an explicitly geothermally-themed collection; 2) improved RESTful (i.e., https-protocol driven) web services for external partners to access xDD data for machine learning application development; 3) a web platform for end-users and xDD administrators to coordinate the development of data mining applications from the initial step of browsing available documents to the final stage of deploying a production-quality machine learning application on high-throughput computing infrastructure; and 4) the development of demonstration data mining applications to illustrate the new workflow to potential collaborators. A total of 21,674 geothermal documents from NGDS were fully ingested into the xDD library and the associated metadata is publicly available through the xDD web services; furthermore, the ADEPT web platform is now publicly accessible and fully live at https://xdd.wisc.edu/adept/.

15 GEOTHERMAL ENERGY↗

Hospice Landscape Report

In July 2018, CMS requested assistance from Oak Ridge National Laboratory (ORNL) to provide expert data science support aimed at developing algorithms for data mining of medical data for operational and payment purposes. The project is intended to be exploratory: work is aimed at alleviating challenges associated with improper payments, specifically audit methodologies and targeting and changes to risk scores. Project goals include developing new, sophisticated methods for audit targeting and improved profile– payment error correlations, specifically focused on Medicare Part C, the program under which MAOs provide health care services to beneficiaries. ORNL conducted RADV analyses against RAPS and EDS data as well as a hospice landscape analysis per a January 2015 dataset that included Medicare beneficiaries who were in hospice in 2017 and 2018. Ongoing work under this project also involves development of predictive models for RADV investigations and hospice landscape.

97 MATHEMATICS AND COMPUTING↗

Structure-aware graph neural network based deep transfer learning framework for enhanced predictive analytics on diverse materials datasets

Abstract Modern data mining methods have demonstrated effectiveness in comprehending and predicting materials properties. An essential component in the process of materials discovery is to know which material(s) will possess desirable properties. For many materials properties, performing experiments and density functional theory computations are costly and time-consuming. Hence, it is challenging to build accurate predictive models for such properties using conventional data mining methods due to the small amount of available data. Here we present a framework for materials property prediction tasks using structure information that leverages graph neural network-based architecture along with deep-transfer-learning techniques to drastically improve the model’s predictive ability on diverse materials (3D/2D, inorganic/organic, computational/experimental) data. We evaluated the proposed framework in cross-property and cross-materials class scenarios using 115 datasets to find that transfer learning models outperform the models trained from scratch in 104 cases, i.e., ≈90%, with additional benefits in performance for extrapolation problems. We believe the proposed framework can be widely useful in accelerating materials discovery in materials science.

Chemistry↗

Thermodynamic properties of the Nd-Bi system via emf measurements, $\mathrm{DFT}$ calculations, machine learning, and $\mathrm{CALPHAD}$ modeling

Thermodynamic properties of the Nd-Bi system were investigated using a combination of experimental measurements, first-principles calculations based on density functional theory (DFT), data mining and machine learning (DM + ML) predictions, and calculation of phase diagrams (CALPHAD) modeling. The electromotive force (emf) of Nd-Bi alloys in molten LiCl-KCl-NdCl 3 at 773–973 K was measured via coulometric titration of Nd into Bi for the determination of thermochemical properties such as activity coefficients and solubilities of Nd in Bi. A new peritectic reaction of [liquid + NdBi 2 = Nd 3 Bi 7 ] at 774 K was confirmed using differential scanning calorimetry, structural (X-ray diffraction), and microstructural (scanning electron microscopy) analyses. The unknown crystal structure of NdBi2 was suggested to be a mixture of the anti-La 2 Sb configuration and the La 2 Te-type configuration based on ML predictions for over 26,000 data-mined AB 2 -type configurations together with DFT-based verifications. Using the newly acquired experimental data and DFT-based calculations, the thermodynamic description of the Nd-Bi system was remodeled, and a more complete Nd-Bi phase diagram was calculated, including the Nd 3 Bi 7 compound, invariant transition reactions, and liquidus temperatures.

36 MATERIALS SCIENCE↗

Multi-range vehicle speed prediction using vehicle connectivity for enhanced energy efficiency of vehicles

An integrated speed prediction framework based on historical traffic data mining and real-time V2I communications for CAVs. The present framework provides multi-horizon speed predictions with different fidelity over short and long horizons. The present multi-horizon speed prediction is integrated with an economic model predictive control (MPC) strategy for the battery thermal management (BTM) of connected and automated electric vehicles (EVs) as a case study. The simulation results over real-world urban driving cycles confirm the enhanced prediction performance of the present data mining strategy over long prediction horizons. Despite the uncertainty in long-range CAV speed predictions, the vehicle level simulation results show that 14% and 19% energy savings can be accumulated sequentially through eco-driving and BTM optimization (eco-cooling), respectively, when compared with normal-driving and conventional BTM strategy.

Amini, Mohammad Reza↗

Landscape composition and configuration have scale-dependent effects on agricultural pest suppression

Increasing landscape heterogeneity (composition and configuration) can enhance natural enemy populations and support pest suppression in agricultural landscapes. Using a network-based data mining approach, we examined independent gradients of landscape composition and configuration at six spatial scales that were associated with pest suppression services measured at 32 sites in Michigan and Wisconsin, USA. We compared the relative effects of landscape composition and configuration across scales with those of local crop type (corn or grassland). We found that multiple gradients of configurational heterogeneity were independent of composition and strongly associated with pest suppression, with different configuration metrics being predictive of pest suppression depending on the spatial scales and regions considered. Landscapes that were more configurationally heterogeneous at smaller spatial scales consistently supported higher pest suppression. In Michigan, pest suppression increased in landscapes with high edge contrast between annual crops and surrounding habitats and high edge density of grassland within 250-500 m radii. In Wisconsin, pest suppression increased with large core area of grassland and high field density within a 250 m radius. The main compositional effect we found was a positive relationship between grassland cover and pest suppression occurring at larger spatial scales (1000-1500 m) and occurring in Wisconsin but not in Michigan. Our findings demonstrate that effects of landscape composition and configuration on pest suppression differ across spatial scales and vary regionally. The network-based data mining techniques used here could be useful for disentangling intercorrelated landscape metrics in a variety of other contexts in landscape ecology.

54 ENVIRONMENTAL SCIENCES↗

Automatic Search of Cataclysmic Variables Based on LightGBM in LAMOST-DR7

The search for special and rare celestial objects has always played an important role in astronomy. Cataclysmic Variables (CVs) are special and rare binary systems with accretion disks. Most CVs are in the quiescent period, and their spectra have the emission lines of Balmer series, HeI, and HeII. A few CVs in the outburst period have the absorption lines of Balmer series. Owing to the scarcity of numbers, expanding the spectral data of CVs is of positive significance for studying the formation of accretion disks and the evolution of binary star system models. At present, the research for astronomical spectra has entered the era of Big Data. The Large Sky Area Multi-Object Fiber Spectroscopy Telescope (LAMOST) has produced more than tens of millions of spectral data. the latest released LAMOST-DR7 includes 10.6 million low-resolution spectral data in 4926 sky regions, providing ideal data support for searching CV candidates. To process and analyze the massive amounts of spectral data, this study employed the Light Gradient Boosting Machine (LightGBM) algorithm, which is based on the ensemble tree model to automatically conduct the search in LAMOST-DR7. Finally, 225 CV candidates were found and four new CV candidates were verified by SIMBAD and published catalogs. This study also built the Gradient Boosting Decision Tree (GBDT), Adaptive Boosting (AdaBoost), and eXtreme Gradient Boosting (XGBoost) models and used Accuracy, Precision, Recall, the F1-score, and the ROC curve to compare the four models comprehensively. Experimental results showed that LightGBM is more efficient. The search for CVs based on LightGBM not only enriches the existing CV spectral library, but also provides a reference for the data mining of other rare celestial objects in massive spectral data.

79 ASTRONOMY AND ASTROPHYSICS↗

Polynomial chaos expansions on principal geodesic Grassmannian submanifolds for surrogate modeling and uncertainty quantification

In this work we introduce a manifold learning-based surrogate modeling framework for uncertainty quantification in high-dimensional stochastic systems. Our first goal is to perform data mining on the available simulation data to identify a set of low-dimensional (latent) descriptors that efficiently parameterize the response of the high-dimensional computational model. To this end, we employ Principal Geodesic Analysis on the Grassmann manifold of the response to identify a set of disjoint principal geodesic submanifolds, of possibly different dimension, that captures the variation in the data. Since operations on the Grassmann require the data to be concentrated, we propose an adaptive algorithm based on Riemannian K-means and the minimization of the sample Fréchet variance on the Grassmann manifold to identify “local” principal geodesic submanifolds that represent different system behavior across the parameter space. Polynomial chaos expansion is then used to construct a mapping between the random input parameters and the projection of the response on these local principal geodesic submanifolds. Here, the method is demonstrated on four test cases, a toy-example that involves points on a hypersphere, a Lotka-Volterra dynamical system, a continuous-flow stirred-tank chemical reactor system, and a two-dimensional Rayleigh-Bénard convection problem.

42 ENGINEERING↗

Nuclear data uncertainty propagation and modeling uncertainty impact evaluation in neutronics core simulation

Uncertainty analysis is a critical requirement in reactor simulation as it is used to quantify the reliability of best-estimate calculation. A comprehensive uncertainty analysis should characterize all sources of uncertainties in a computationally-feasible and scientifically-defendable manner. Here we employ a well-established reduced order modeling (ROM) based uncertainty quantification methodology to propagate uncertainties throughout neutronic calculations. ROM relies on recent advances in randomized data mining techniques applied to large data streams. In our proposed implementation, the nuclear data uncertainties are first propagated from multi-group level through lattice physics calculation to generate few-group parameter uncertainties, described using a vector of mean values and a covariance matrix. Employing an ROM-based compression of the covariance matrix, the few-group uncertainties are then propagated through downstream core simulation in a computationally efficient manner. This straightforward approach, albeit efficient as compared to brute force forward and/or adjoint-based methods, often employs a number of assumptions that have been unquestioned in the literature of neutronic uncertainty analysis. This manuscript argues that these assumptions could introduce another source of uncertainty referred to as modeling uncertainties, whose magnitude needs to be quantified in tandem with nuclear data uncertainties. Thus, our primary goal is to explore the interactions between these two uncertainty sources in order to assess whether modeling uncertainties have an impact on parameter uncertainties. To explore this endeavor, the impact of a number of modeling assumptions on core attributes uncertainties is quantified. The study employs a CANDU reactor model, with Serpent and NEWT as lattice physics solvers and NESTLE-C as core simulator. The modeling assumptions investigated include those related with the uncertainty propagation method employed, e.g., deterministic vs. stochastic, the few-group energy structure employed to represent the cross-sections, the resonance treatment in lattice physics calculation, the reference values for the cross-section, and the number of samples employed to render ROM compression. Results indicate that some of the modeling assumptions could have a non-negligible impact on the core responses propagated uncertainties, highlighting the need for a more comprehensive approach to combine parameter and modeling uncertainties.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Addressing Data Center Cooling Needs through the Use of Subsurface Thermal Energy Storage Systems

This study aims to evaluate the feasibility of addressing the cooling needs for information technology (IT) equipment in data centers by using reservoir thermal energy storage (RTES) to provide reliable and sustainable low-temperature fluid. This project focuses on the technical viability of operating such a system in Houston, Texas, which is a representative location for crypto mining data centers. An analysis has been performed to investigate the technical feasibility with climate data for Houston. Results show that data centers on the scale of 30 MW, and operating at a temperature of 27 degrees C, can be reliably cooled by a combined RTES and dry coolers setup. A techno-economic analysis will be performed and energy/water saving benefits will be quantified in the future. In addition, the study will be extended to other representative data center locations with different climates and geographical locations.

cooling↗

HPRT promotes proliferation and metastasis in head and neck squamous cell carcinoma through direct interaction with STAT3

Highlights: • The higher expression of HPRT in HNSCC tissue was related to poor prognosis of patients. • Overexpression of HPRT increased the gene expression of epithelial mesenchymal transition markers. • HPRT directly interacted with STAT3. • Knocking down HPRT significantly decreased tumour growth and enhanced the anticancer effect against HNSCC xenografts. Increasing effort has been put into finding novel molecular pathways to improve the efficiency of EGFR inhibitors against head and neck squamous cell cancer (HNSCC). In this study, we performed data mining and bioinformatically analysed RNA-Seq data downloaded from TCGA and confirmed that higher expression of HPRT in HNSCC tissue was related to poor prognosis of patients. Then, we conducted in vitro and in vivo loss- and gain-of-function experiments to demonstrate the role of HPRT in HNSCC cell lines. Overexpression of HPRT increased the gene expression of epithelial mesenchymal transition markers via direct interaction with STAT3. Knocking down HPRT significantly decreased tumour growth and enhanced the anticancer effect of EGFR inhibitors against HNSCC xenografts. In conclusion, HPRT is a binding partner of STAT3 that promotes EMT and proliferation. Our findings support HPRT as a promising prognostic indicator and potential therapeutic target for HNSCC.

60 APPLIED LIFE SCIENCES↗

LIB Design Module for Grid Energy System Application

We will employ a machine learning approach with intelligent data mining and database construction to analyze enormous data repositories for identifying and extracting geographic-dependent cell design specifications from publicly accessible grid-scale energy storage usage databases in an automated way at scale.

Liu, Dianying [Pacific Northwest National Laborato↗