Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data mining”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

Putting Priors in Mixture Density Mercer Kernels

This paper presents a new methodology for automatic knowledge driven data mining based on the theory of Mercer Kernels, which are highly nonlinear symmetric positive definite mappings from the original image space to a very high, possibly infinite dimensional feature space. We describe a new method called Mixture Density Mercer Kernels to learn kernel function directly from data, rather than using predefined kernels. These data adaptive kernels can en- code prior knowledge in the kernel using a Bayesian formulation, thus allowing for physical information to be encoded in the model. We compare the results with existing algorithms on data from the Sloan Digital Sky Survey (SDSS). The code for these experiments has been generated with the AUTOBAYES tool, which automatically generates efficient and documented C/C++ code from abstract statistical model specifications. The core of the system is a schema library which contains template for learning and knowledge discovery algorithms like different versions of EM, or numeric optimization methods like conjugate gradient methods. The template instantiation is supported by symbolic- algebraic computations, which allows AUTOBAYES to find closed-form solutions and, where possible, to integrate them into the code. The results show that the Mixture Density Mercer-Kernel described here outperforms tree-based classification in distinguishing high-redshift galaxies from low- redshift galaxies by approximately 16% on test data, bagged trees by approximately 7%, and bagged trees built on a much larger sample of data by approximately 2%.

Srivastava, Ashok N.↗

An Integrated Approach for Gear Health Prognostics

In this paper, an integrated approach for gear health prognostics using particle filters is presented. The presented method effectively addresses the issues in applying particle filters to gear health prognostics by integrating several new components into a particle filter: (1) data mining based techniques to effectively define the degradation state transition and measurement functions using a one-dimensional health index obtained by whitening transform; (2) an unbiased l-step ahead RUL estimator updated with measurement errors. The feasibility of the presented prognostics method is validated using data from a spiral bevel gear case study.

He, David↗

Transcriptome Mining Provides Insights into Cell Wall Metabolism and Fiber Lignification in Agave tequilana Weber

Resilience of growing in arid and semiarid regions and a high capacity of accumulating sugar-rich biomass with low lignin percentages have placed Agave species as an emerging bioenergy crop. Although transcriptome sequencing of fiber-producing agave species has been explored, molecular bases that control wall cell biogenesis and metabolism in agave species are still poorly understood. Here, through RNAseq data mining, we reconstructed the cellulose biosynthesis pathway and the phenylpropanoid route producing lignin monomers in A. tequilana, and evaluated their expression patterns in silico and experimentally. Most of the orthologs retrieved showed differential expression levels when they were analyzed in different tissues with contrasting cellulose and lignin accumulation. Phylogenetic and structural motif analyses of putative CESA and CAD proteins allowed to identify those potentially involved with secondary cell wall formation. RT-qPCR assays revealed enhanced expression levels of AtqCAD5 and AtqCESA7 in parenchyma cells associated with extraxylary fibers, suggesting a mechanism of formation of sclerenchyma fibers in Agave similar to that reported for xylem cells in model eudicots. Overall, our results provide a framework for understanding molecular bases underlying cell wall biogenesis in Agave species studying mechanisms involving in leaf fiber development in monocots.

59 BASIC BIOLOGICAL SCIENCES↗

Software Tools Streamline Project Management

Three innovative software inventions from Ames Research Center (NETMARK, Program Management Tool, and Query-Based Document Management) are finding their way into NASA missions as well as industry applications. The first, NETMARK, is a program that enables integrated searching of data stored in a variety of databases and documents, meaning that users no longer have to look in several places for related information. NETMARK allows users to search and query information across all of these sources in one step. This cross-cutting capability in information analysis has exponentially reduced the amount of time needed to mine data from days or weeks to mere seconds. NETMARK has been used widely throughout NASA, enabling this automatic integration of information across many documents and databases. NASA projects that use NETMARK include the internal reporting system and project performance dashboard, Erasmus, NASA s enterprise management tool, which enhances organizational collaboration and information sharing through document routing and review; the Integrated Financial Management Program; International Space Station Knowledge Management; Mishap and Anomaly Information Reporting System; and management of the Mars Exploration Rovers. Approximately $1 billion worth of NASA s projects are currently managed using Program Management Tool (PMT), which is based on NETMARK. PMT is a comprehensive, Web-enabled application tool used to assist program and project managers within NASA enterprises in monitoring, disseminating, and tracking the progress of program and project milestones and other relevant resources. The PMT consists of an integrated knowledge repository built upon advanced enterprise-wide database integration techniques and the latest Web-enabled technologies. The current system is in a pilot operational mode allowing users to automatically manage, track, define, update, and view customizable milestone objectives and goals. The third software invention, Query-Based Document Management (QBDM) is a tool that enables content or context searches, either simple or hierarchical, across a variety of databases. The system enables users to specify notification subscriptions where they associate "contexts of interest" and "events of interest" to one or more documents or collection(s) of documents. Based on these subscriptions, users receive notification when the events of interest occur within the contexts of interest for associated document or collection(s) of documents. Users can also associate at least one notification time as part of the notification subscription, with at least one option for the time period of notifications.

Source record↗

Robotic Rock Classification

This report describes a three-month research program undertook jointly by the Robotics Institute at Carnegie Mellon University and Ames Research Center as part of the Ames' Joint Research Initiative (JRI.) The work was conducted at the Ames Research Center by Mr. Liam Pedersen, a graduate student in the CMU Ph.D. program in Robotics under the supervision Dr. Ted Roush at the Space Science Division of the Ames Research Center from May 15 1999 to August 15, 1999. Dr. Martial Hebert is Mr. Pedersen's research adviser at CMU and is Principal Investigator of this Grant. The goal of this project is to investigate and implement methods suitable for a robotic rover to autonomously identify rocks and minerals in its vicinity, and to statistically characterize the local geological environment. Although primary sensors for these tasks are a reflection spectrometer and color camera, the goal is to create a framework under which data from multiple sensors, and multiple readings on the same object, can be combined in a principled manner. Furthermore, it is envisioned that knowledge of the local area, either a priori or gathered by the robot, will be used to improve classification accuracy. The key results obtained during this project are: The continuation of the development of a rock classifier; development of theoretical statistical methods; development of methods for evaluating and selecting sensors; and experimentation with data mining techniques on the Ames spectral library. The results of this work are being applied at CMU, in particular in the context of the Winter 99 Antarctica expedition in which the classification techniques will be used on the Nomad robot. Conversely, the software developed based on those techniques will continue to be made available to NASA Ames and the data collected from the Nomad experiments will also be made available.

Hebert, Martial↗

Automated Knowledge Discovery From Simulators

A computational method, SimLearn, has been devised to facilitate efficient knowledge discovery from simulators. Simulators are complex computer programs used in science and engineering to model diverse phenomena such as fluid flow, gravitational interactions, coupled mechanical systems, and nuclear, chemical, and biological processes. SimLearn uses active-learning techniques to efficiently address the "landscape characterization problem." In particular, SimLearn tries to determine which regions in "input space" lead to a given output from the simulator, where "input space" refers to an abstraction of all the variables going into the simulator, e.g., initial conditions, parameters, and interaction equations. Landscape characterization can be viewed as an attempt to invert the forward mapping of the simulator and recover the inputs that produce a particular output. Given that a single simulation run can take days or weeks to complete even on a large computing cluster, SimLearn attempts to reduce costs by reducing the number of simulations needed to effect discoveries. Unlike conventional data-mining methods that are applied to static predefined datasets, SimLearn involves an iterative process in which a most informative dataset is constructed dynamically by using the simulator as an oracle. On each iteration, the algorithm models the knowledge it has gained through previous simulation trials and then chooses which simulation trials to run next. Running these trials through the simulator produces new data in the form of input-output pairs. The overall process is embodied in an algorithm that combines support vector machines (SVMs) with active learning. SVMs use learning from examples (the examples are the input-output pairs generated by running the simulator) and a principle called maximum margin to derive predictors that generalize well to new inputs. In SimLearn, the SVM plays the role of modeling the knowledge that has been gained through previous simulation trials. Active learning is used to determine which new input points would be most informative if their output were known. The selected input points are run through the simulator to generate new information that can be used to refine the SVM. The process is then repeated. SimLearn carefully balances exploration (semi-randomly searching around the input space) versus exploitation (using the current state of knowledge to conduct a tightly focused search). During each iteration, SimLearn uses not one, but an ensemble of SVMs. Each SVM in the ensemble is characterized by different hyper-parameters that control various aspects of the learned predictor - for example, whether the predictor is constrained to be very smooth (nearby points in input space lead to similar output predictions) or whether the predictor is allowed to be "bumpy." The various SVMs will have different preferences about which input points they would like to run through the simulator next. SimLearn includes a formal mechanism for balancing the ensemble SVM preferences so that a single choice can be made for the next set of trials.

Burl, Michael↗

Space Motion Sickness - Analysis of Medical Debriefs Data for Incidence and Treatment

Astronauts use medications for the treatment of a variety of illnesses during space travel. Data mining efforts to assess minor clinical conditions occurring during Shuttle flights STS-1 through STS-94 revealed that space motion sickness (SMS) was the most common ailment during early flight days, occurring in approx.40% of crewmembers, followed by digestive system disturbances (9%) and infectious diseases, which most commonly involved the respiratory or urinary tracts. A more recent analysis of postflight medical debriefs data to examine trends with respect to medication use by astronauts during spaceflights indicated that ~37% of all prescriptions recorded was for pain followed by sleep (22%), SMS (18%), decongestion (14%), and all others (14%). Further analysis revealed that about 150 of 317 crewmembers experienced symptoms of SMS. Nearly all (132 of 150) crewmembers took medication for the treatment of symptoms with a total of 387 doses. Promethazine was taken most often (201 doses); in most cases this resulted in alleviation of symptoms with 130 crewmembers (65%) reporting feeling much or somewhat better. Although fewer total doses of the combination of promethazine and dextroamphetamine (Phen/Dex) were taken (45 doses), slightly more than half of these doses resulted in improvement. The combination of scopolamine and dextroamphetamine (Scop/Dex) was reported to be effective in only 37% of cases, with 36 of 97 total doses resulting in improvement. A higher percentage (24%) of Scop/Dex doses was reported to be ineffective compared with promethazine alone or as Phen/Dex (10% and 7%, respectively). Comparisons of the effectiveness of the different dosage forms of promethazine revealed that intramuscular injection was most effective in alleviating symptoms with 55% feeling much better, 16% feeling somewhat better, and only 7% feeling no effect or worse. Overall, it appears that promethazine alone was used more frequently during flight and was reported effective for the treatment of SMS.

Putcha, Lakshmi↗

System Engineers and Decisions: It?s All about Knowledge

In order to guarantee that a system meets adequate levels of reliability and availability, system performances are continuously monitored and analyzed thanks to the technological advancements driving the Industry 4.0 revolution. An Industry 4.0 approach is typically based on advanced statistical, big data mining, machine learning, and internet-of-things methods designed to detect anomalies in the behavior of system, detect the most likely failure modes, and provide indications to system engineers on when maintenance activities should be performed before system performance are deemed unacceptable (which can be generated by diagnostic and prognostic methods). However, these analyses, which are designed to automatize and increase the efficacy of the system maintenance program, require large amount of data which can come in various forms: numeric, textual, images, sounds etc. Such data constitutes the historic knowledge benchmark to track system performances and support system engineer decisions. Here we claim that data is not sufficient to support this kind of analyses when applied to systems characterized by complex architectures and behaviors. Robust system engineer decisions require the ability to understand the system operational context that lies behind the observed data elements. In this respect, system models are in fact necessary to “put data in context” and capture relationships between data elements. Industry 4.0 methods require in fact contextual knowledge as a basis upon which hypotheses can be generated and assumptions tested. In our view, for complex systems, model-based system engineering (MBSE) models can afford this contextual knowledge, as they are typically used to describe systems architecture and dynamic behaviors. System knowledge is here intended as the blending of collected data and system architecture which takes the form of a “knowledge graph”. A knowledge graph is a database which consists of a large set of nodes (in our case an entity can be either a data or an MBSE element) which are linked to each other. The types of nodes and links follow a pre-defined topology, sometimes also refers as an ontology, that is designed to fit the actual decisions that needs to be performed. We show here how a knowledge graph can be defined to support system engineer maintenance decisions and how the same graph can be built based on system MBSE models and pre-processed data from numeric (through anomaly detections and diagnostic methods) and textual elements (through technical language processing TLP).

97 - MATHEMATICS AND COMPUTING↗

NASA Exhibits

A series of NASA presentations for the Supercomputing 2001 conference are summarized. The topics include: (1) Mars Surveyor Landing Sites "Collaboratory"; (2) Parallel and Distributed CFD for Unsteady Flows with Moving Overset Grids; (3) IP Multicast for Seamless Support of Remote Science; (4) Consolidated Supercomputing Management Office; (5) Growler: A Component-Based Framework for Distributed/Collaborative Scientific Visualization and Computational Steering; (6) Data Mining on the Information Power Grid (IPG); (7) Debugging on the IPG; (8) Debakey Heart Assist Device: (9) Unsteady Turbopump for Reusable Launch Vehicle; (10) Exploratory Computing Environments Component Framework; (11) OVERSET Computational Fluid Dynamics Tools; (12) Control and Observation in Distributed Environments; (13) Multi-Level Parallelism Scaling on NASA's Origin 1024 CPU System; (14) Computing, Information, & Communications Technology; (15) NAS Grid Benchmarks; (16) IPG: A Large-Scale Distributed Computing and Data Management System; and (17) ILab: Parameter Study Creation and Submission on the IPG.

Deardorff, Glenn↗

Automated annotation of scientific texts for ML-based keyphrase extraction and validation

Advanced omics technologies and facilities generate a wealth of valuable data daily; however, the data often lack the essential metadata required for researchers to find, curate, and search them effectively. The lack of metadata poses a significant challenge in the utilization of these data sets. Machine learning (ML)–based metadata extraction techniques have emerged as a potentially viable approach to automatically annotating scientific data sets with the metadata necessary for enabling effective search. Text labeling, usually performed manually, plays a crucial role in validating machine-extracted metadata. However, manual labeling is time-consuming and not always feasible; thus, there is a need to develop automated text labeling techniques in order to accelerate the process of scientific innovation. This need is particularly urgent in fields such as environmental genomics and microbiome science, which have historically received less attention in terms of metadata curation and creation of gold-standard text mining data sets. In this paper, we present two novel automated text labeling approaches for the validation of ML-generated metadata for unlabeled texts, with specific applications in environmental genomics. Our techniques show the potential of two new ways to leverage existing information that is only available for select documents within a corpus to validate ML models, which can then be used to describe the remaining documents in the corpus. The first technique exploits relationships between different types of data sources related to the same research study, such as publications and proposals. The second technique takes advantage of domain-specific controlled vocabularies or ontologies. In this paper, we detail applying these approaches in the context of environmental genomics research for ML-generated metadata validation. Our results show that the proposed label assignment approaches can generate both generic and highly specific text labels for the unlabeled texts, with up to 44% of the labels matching with those suggested by a ML keyword extraction algorithm.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Hybrid Power Plants: Status of Operating and Proposed Plants, 2022 Edition [Slides]

Falling battery prices and the growth of variable renewable generation are driving a surge of interest in “hybrid” power plants that combine, for example, wind or solar generating capacity with co-located batteries. While most of the current interest involves pairing photovoltaic (PV) plants with batteries, other types of hybrid or co-located plants with wide-ranging configurations have been part of the U.S. electricity mix for decades. This annually updated briefing tracks and maps existing hybrid or co-located plants across the United States while also synthesizing data mined from power purchase agreements (PPAs) and generation interconnection queues to shed light on near- and long-term development pipelines. The scope includes co-located hybrid plants that pair two or more generators and/or that pair generation with storage at a single point of interconnection, and full hybrids that feature co-location and co-control. The focus is on plants with one megawatt (MW) or more of capacity; smaller (often behind-the-meter) projects are also increasingly common, but are not included in this data synthesis. Key findings from the latest briefing include: -At the end of 2021, there were nearly 300 hybrid plants (>1 MW) operating across the United States, totaling nearly 36 gigawatts (GW) of generating capacity and 3.2 GW/8.1 GWh of energy storage. PV+storage plants are by far the most common, dominating in terms of plant number (140), storage capacity (2.2 GW/7.0 GWh), storage:generator ratio (53%), and storage duration (3.2 hours). But there are nearly twenty other hybrid plant configurations as well, including several different fossil hybrid categories (each dominated by the fossil component) as well as wind+storage, wind+PV, wind+PV+storage, geothermal+PV, and others. -Last year was a breakout year for PV+storage hybrids in particular: 67 of the 74 hybrids added in 2021 were PV+storage. By the end of 2021, there were more GW of battery capacity installed in PV+storage hybrids (2.2 GW) than as standalone storage plants (1.8 GW). The difference is even starker in energy terms, with PV+storage plants hosting twice as much battery capacity as standalone storage plants (7 GWh vs. 3.5 GWh, respectively). Much of the battery capacity added in hybrid form in 2021 was a battery retrofit to a pre-existing PV plant. -Data on plants under development from the interconnection queues of all seven ISOs/RTOs plus 35 individual utilities suggest that these hybridization trends are likely to continue. At the close of 2021, there were more than 670 GW of solar plants in the nation’s queues; 285 GW (~42%) of this capacity was proposed as a hybrid, most typically pairing PV with battery storage (PV+storage represented nearly 90% of all hybrid capacity in the queues). For wind, 247 GW of capacity sat in the queues, with 19 GW (~8%) proposed as a hybrid, again most-often pairing wind with storage (wind+storage represented ~4% of all hybrid capacity in the queues). Meanwhile, nearly half of all storage in the queues is estimated to be part of a hybrid plant. While many of these proposed plants will not ultimately reach commercial operations, the depth of interest in hybrid plants—especially PV+storage—is notable.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Spacesuit Glove-Induced Hand Trauma and Analysis of Potentially Related Risk Variables

Injuries to the hands are common among astronauts who train for extravehicular activity (EVA). When the gloves are pressurized, they restrict movement and create pressure points during tasks, sometimes resulting in pain, muscle fatigue, abrasions, and occasionally more severe injuries such as onycholysis. Glove injuries, both anecdotal and recorded, have been reported during EVA training and flight persistently through NASA's history regardless of mission or glove model. Theories as to causation such as glove-hand fit are common but often lacking in supporting evidence. Previous statistical analysis has evaluated onycholysis in the context of crew anthropometry only (Opperman et al 2010). The purpose of this study was to analyze all injuries (as documented in the medical records) and available risk factor variables with the goal to determine engineering and operational controls that may reduce hand injuries due to the EVA glove in the future. A literature review and data mining study were conducted between 2012 and 2014. This study included 179 US NASA crew who trained or completed an EVA between 1981 and 2010 (crossing both Shuttle and ISS eras) and wore either the 4000 Series or Phase VI glove during Extravehicular Mobility Unit (EMU) spacesuit EVA training and flight. All injuries recorded in medical records were analyzed in their association to candidate risk factor variables. Those risk factor variables included demographic characteristics, hand anthropometry, glove fit characteristics, and training/EVA characteristics. Utilizing literature, medical records and anecdotal causation comments recorded in crewmember injury data, investigators were able to identify several risk factors associated with increased risk of glove related injuries. Prime among them were smaller hand anthropometry, duration of individual suited exposures, and improper glove-hand fit as calculated by the difference in the anthropometry middle finger length compared to the baseline EVA glove middle finger length.

McFarland, Shane M.↗

Spacesuit Glove-Induced Hand Trauma and Analysis of Potentially Related Risk Variables

Injuries to the hands are common among astronauts who train for extravehicular activity (EVA). When the gloves are pressurized, they restrict movement and create pressure points during tasks, sometimes resulting in pain, muscle fatigue, abrasions, and occasionally more severe injuries such as onycholysis. Glove injuries, both anecdotal and recorded, have been reported during EVA training and flight persistently through NASA's history regardless of mission or glove model. Theories as to causation such as glove-hand fit are common but often lacking in supporting evidence. Previous statistical analysis has evaluated onycholysis in the context of crew anthropometry only. The purpose of this study was to analyze all injuries (as documented in the medical records) and available risk factor variables with the goal to determine engineering and operational controls that may reduce hand injuries due to the EVA glove in the future. A literature review and data mining study were conducted between 2012 and 2014. This study included 179 US NASA crew who trained or completed an EVA between 1981 and 2010 (crossing both Shuttle and ISS eras) and wore either the 4000 Series or Phase VI glove during Extravehicular Mobility Unit (EMU) spacesuit EVA training and flight. All injuries recorded in medical records were analyzed in their association to candidate risk factor variables. Those risk factor variables included demographic characteristics, hand anthropometry, glove fit characteristics, and training/EVA characteristics. Utilizing literature, medical records and anecdotal causation comments recorded in crewmember injury data, investigators were able to identify several risk factors associated with increased risk of glove related injuries. Prime among them were smaller hand anthropometry, duration of individual suited exposures, and improper glove-hand fit as calculated by the difference in the anthropometry middle finger length compared to the baseline EVA glove middle finger length.

Charvat, Chacqueline M.↗

Distributed non-negative matrix factorization with determination of the number of latent features

The holistic analysis and understanding of the latent (that is, not directly observable) variables and patterns buried in large datasets is crucial for data-driven science, decision making and emergency response. Such exploratory analyses require devising unsupervised learning methods for data mining and extraction of the latent features, and non-negative matrix factorization (NMF) is one of the prominent such methods. NMF is based on compute-intense non-convex constrained minimization, which, for large datasets requires fast and distributed algorithms. However, current parallel implementations of NMF fail to estimate the number of latent features. In practice, identifying these features is both difficult and significant for pattern recognition and latent feature analysis, especially for large dense matrices. Here, we introduce a distributed NMF algorithm coupled with distributed custom clustering followed by a stability analysis on dense data, which we call DnMFk, to determine the number of latent variables. The results on synthetic data and the classical Swimmer data set demonstrate the accuracy of model determination while scaling nearly linearly across multiple processors for large data. Further, we employ DnMFk to determine the number of hidden features from a terabyte matrix.

97 MATHEMATICS AND COMPUTING↗

Physics of Boundaries and their Interactions in Space Plasmas

This final report describes a brief summary of our accomplishments during the complete contract period. Traditionally, due to computational limitations, it has been impossible to obtain a global view of the magnetosphere on ion time and spatial scales. As a result, kinetic simulations have concentrated on the local structure of different magnetospheric discontinuities and boundaries. However, due to the emergence of low cost desktop superconductors, as well as by taking full advantage of latest advances in data mining and visualization technology, we were able to bypass our planned (proposed) regional simulations and proceed to large-scale 3-D and 2-D global hybrid simulations of the magnetosphere. As a result, although we are only finishing the second year of the proposed activity, much of the original scientific objectives have been surpassed and new avenues of investigation have been opened. Such simulations have led us to possible explanations of some long-standing issues in magnetospheric physics. They have also enabled us to make a number of important discoveries/predictions, which need to be looked for in satellite data. Examples include: (1) the finding that the bow shock can become unstable to the Kelvin-Helmholtz (KH;) (2) the discovery of a mechanism for intermittent reconnection due to ion physics which may be relevant to the explanation of the recurrence rate of flux transfer events (FTEs;) and (3) the finding that the current sheet in the near-Earth magnetotail region can become unstable to KH with detectable, unique ionospheric signatures. Further, we demonstrated a viable mechanism for the onset of reconnection at the magnetopause, examined the detailed structure of the boundary layer incorporating curvature effects, and provided an explanation for the large core fields observed within FTEs as well as flux ropes in the magnetotail.

Omidi, Nojan↗

Physics of Boundaries and their Interactions in Space Plasmas

This final report describes a brief summary of our accomplishments during the complete contract period. Traditionally, due to computational limitations, it has been impossible to obtain a global view of the magnetosphere on ion time and spatial scales. As a result, kinetic-simulations have concentrated on the local structure of different magnetospheric discontinuities and boundaries. However, due to the emergence of low cost supercomputers, as well as by taking full advantage of latest advances in data mining and visualization technology, we were able to bypass our planned (proposed) regional simulations and proceed to large-scale 3-D and 2-D global hybrid simulations of the magnetosphere. As a result, although we are only finishing the second year of the proposed activity, much of the original scientific objectives have been surpassed and new avenues of investigation have been opened. Such simulations have led us to possible explanations of some long-standing issues in magnetospheric physics. They have also enables us to make a number of important discoveries predictions, which need to be looked for in satellite data. Examples include the finding that the bow shock can become unstable to the Kelvin-Helmholtz (KH), (2) the discovery of a mechanism for intermittent reconnection due to ion physics which may be relevant to the explanation of the recurrence rate of flux transfer events (FTEs), and (3) this finding that the current sheet in the near-Earth magnetotail region can become unstable to KH with detectable, unique ionospheric signatures. Further, we demonstrated a viable mechanism for the onset of reconnection at the magnetopause, examined the detailed structure of the boundary layer incorporating curvature effects, and provided an explanation for the large core fields observed within FTEs as well as flux ropes in the magnetotail.

Omidi, Nojan↗

CORE: A Global Aggregation Service for Open Access Papers

This paper introduces CORE, a widely used scholarly service, which provides access to the world’s largest collection of open access research publications, acquired from a global network of repositories and journals. CORE was created with the goal of enabling text and data mining of scientific literature and thus supporting scientific discovery, but it is now used in a wide range of use cases within higher education, industry, not-for-profit organisations, as well as by the general public. Through the provided services, CORE powers innovative use cases, such as plagiarism detection, in market-leading third-party organisations. CORE has played a pivotal role in the global move towards universal open access by making scientific knowledge more easily and freely discoverable. In this paper, we describe CORE’s continuously growing dataset and the motivation behind its creation, present the challenges associated with systematically gathering research papers from thousands of data providers worldwide at scale, and introduce the novel solutions that were developed to overcome these challenges. The paper then provides an in-depth discussion of the services and tools built on top of the aggregated data and finally examines several use cases that have leveraged the CORE dataset and services.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Predicting Execution Times for Disk-based and In-Situ Parallel Data Analytics (Final Technical Report)

In recent years, there has been a significant amount of interests in in-situ analytics on simulation programs. For a variety of reasons, it is desirable to be able to predict the execution time of an analytics program. At the same time, frameworks such as MapReduce have become popular for scientific data analytics. This paper focuses on developing performance models for predicting execution time of parallel data analytics, with a special emphasis on in-situ analytics. We take two distinct approach towards performance prediction. We first expand SKOPE (a SKeleton framewOrk for Performance Exploration) with performance models for disk data read, cache performance, and page fault penalty. Second, an analytical performance model is also developed. We have evaluated our performance prediction framework as well as the analytical model on three hardware setups with well-known data mining algorithms implemented in three programming paradigms, MapReduce, MATE (a MapReduce-like parallel system with an alternate API for multi-core environments) and Smart (a MapReduce-like framework for in-situ analytics). Results show that our performance prediction framework along with the incorporated performance models are capable of accurately predicting execution times for parallel scientific analytics on different hardware setups.

97 MATHEMATICS AND COMPUTING↗