Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “imprecise labels”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

Root identification in minirhizotron imagery with multiple instance learning

In this study, multiple instance learning (MIL) algorithms to automatically perform root detection and segmentation in minirhizotron imagery using only image-level labels are proposed. Root and soil characteristics vary from location to location, and thus, supervised machine learning approaches that are trained with local data provide the best ability to identify and segment roots in minirhizotron imagery. However, labeling roots for training data (or otherwise) is an extremely tedious and time-consuming task. This paper aims to address this problem by labeling data at the image level (rather than the individual root or root pixel level) and train algorithms to perform individual root pixel level segmentation using MIL strategies. Three MIL methods (multiple instance adaptive cosine coherence estimator, multiple instance support vector machine, multiple instance learning with randomized trees) were applied to root detection and compared to non-MIL approaches. The results show that MIL methods improve root segmentation in challenging minirhizotron imagery and reduce the labeling burden. In our results, multiple instance support vector machine outperformed other methods. The multiple instance adaptive cosine coherence estimator algorithm was a close second with an added advantage that it learned an interpretable root signature which identified the traits used to distinguish roots from soil and did not require parameter selection.

59 BASIC BIOLOGICAL SCIENCES↗

Use of Machine Learning on PMU Data for Transmission System Fault Analysis

Synchrophasor technology has been used for monitoring, control, and protection of bulk power system for over 10 years. Deployment of phasor measurement units (PMUs) in the USA power system has surpassed 3000 units installed in the transmission substations as stand-alone intelligent electronic devices (IEDs) or as a software add-on to other devices such as digital protective relays (DPRs) or digital fault recorders (DFRs). By now, thousands of terabytes of PMU data may have been captured and stored by various transmission system operators (TSOs) and independent system operators (ISOs). This creates an opportunity to deploy advanced machine learning (ML) techniques to detect and classify faults recorded by PMUs automatically to be used by the system operators for rapid, critical decision-making when manual analysis of the past or unfolding events is not feasible. In this paper we offer a brief background on how the automated fault analysis may be done using DPR and/or DFR data, and compare some of the legacy approaches to the new ML approaches in the context of the system-wide PMU recordings. We then offer insights from developing practical ML solutions that have been applied on field recordings captured by close to 450 PMUs from all three US interconnections (Western, Eastern and ERCOT) over two years (2016-2017). We identify and illustrate ML challenges we addressed: inaccurate data, data with scarce and temporally imprecise fault labels, data recorded by PMUs sparsely located at substations resulting in the fault records taken afar from the ends of the faulted lines, data containing only positive sequence values, and data taken at different voltage levels. We then illustrate the ML model results for fault analysis under different application scenarios. The novelty of this study is not only in the design, implementation, and performance analysis of the ML algorithms, but also in the use of advanced fault modelling and simulation approaches to improve the training results when developing supervised ML models for fault detection and classification. Extensive simulations of faults were conducted on a 14-bus power system to create a training dataset with over 1400 accurately labelled faults. This dataset was applied to enhance the accuracy of fault detection and classification of machine learning-based models trained with small number of labelled faults in large datasets recorded in the grid interconnections ranging from 5,000 to 70,000 buses.

Synchrophasors, Machine Learning, Fault Analysis, ↗

Rays for Roots - Integrating Backscatter X-Ray Phenotyping, Modeling and Genetics to Increase Carbon Sequestration and Switchgrass Resource Use (Final Report)

To increase carbon (C) deposition in the soil and enhance crop resource use efficiency, characterizing root form and function is essential. Several root and soil traits have been linked to increased root-to-soil C transfer. Technology that could provide high-resolution characterization of many of these traits in field conditions would revolutionize our ability to study and understand how to increase C sequestration. In this effort, we developed an initial early prototype backscatter X-ray system for non-destructive imaging of root traits. We collected initial backscatter X-ray data in field and lab settings and carried out early analysis of these data. Along with this prototype, we also developed a suite of root phenotyping approaches including advanced minirhizotron image analysis, soil core imaging, and mesocosm imaging. Minirhizotron (MR) tubes are clear tubes inserted into the soil in the field and used to image roots and the surrounding soil. Our team has developed deep learning-based methods that can segment roots from soil that can learn from imprecise image-level labels. The ability to learn or fine-tune our deep learning algorithms from image-level labels allows easier and faster application of these approaches to new locations and new plant species. We have successfully implemented and applied our MR analysis approaches to thousands of switchgrass MR images collected across geographical regions. An advantage of MR imaging is the ability to collect root and soil images over time. Our soil core analysis included collecting hundreds of soil core samples from harvested switchgrass fields and imaging these cores with both X-ray CT and backscatter X-ray imaging. Initial segmentation approaches for the X-ray CT images of these cores have been developed and applied. An advantage of soil core analysis is that it preserves the three-dimensional structures of the roots and soil in the core collected. Our group also developed photogrammetry-based mesocosm root imaging and phenotyping approaches. In this approach, a plant was grown in a large mesocosm with a three-dimensional grid of thin supporting lines inserted throughout the mesocosm. After the plant (and, correspondingly, the root architecture is grown and established) the soil media was removed and the supporting lines approximately preserved the three-dimensional root architecture. Then, we applied photogrammetry techniques to create a three-dimensional digital representation of the root architecture for which we developed analysis algorithms including skeletonization. We carried out our phenotyping development with powerful switchgrass resources and physiological and agroecosystem modeling to deliver novel technology. This project contributes to multiple ARPA-E missions including reduction of foreign imports of energy, reduction of energy-related emissions including greenhouse gases, and ensuring that the United States maintains a technological lead in developing and deploying advanced energy technology. Furthermore, the developed tools could transform public and private plant breeding and could be broadly applicable to other crops and, potentially, other application areas. Our team of engineers, plant and soil scientists, and modelers i) developed an early prototype backscatter X-ray platform that can operate in field conditions; ii) developed a suite of root phenotyping and characterization approaches as described above; iii) developed and carried out plant biology and physiology roots studies and; iv) developed and implemented mechanistic physiological modeling.

42 ENGINEERING↗

Big Data Synchrophasor Monitoring and Analytics for Resiliency Tracking (BDSMART)

This report contains key findings from a project titled Big Data Synchrophasor Monitoring and Analytics for Resiliency Tracking (BDSMART), which was carried out through a collaborative effort of a team of researchers from Texas A&M Engineering Experiment Station, Temple University, and Quanta Technology, LLC. The in-kind support came from OSIsoft (acquired by AVEVA), which provided their PI Historian software to demonstrate the use case of streaming PMU data. The first section of the report describes the project goals and objectives related to the development of Machine Learning (ML) models capable of detecting and classifying events by processing phasor measurements captured in the field by Phasor Measurement Units (PMUs). The data for this study was contributed by the utilities/ISOs from the Western and Eastern interconnects and ERCOT, further referred to as Interconnect B (IC B), Interconnect A (IC A), and Interconnect C (IC C), respectively. The approach that the BDSMART Research Team proposed and the key research tasks defined by the team are outlined in this section. The next section describes the technical approach. We first discuss the data constraints related to the PMU measurements and data interpretation constraints imposed by the data contributors. They provided neither the topological information of the grid nor PMU placement locations and captured recorded data at very few locations in the system with the reporting rate of either 30 or 60 fps. The recordings are mostly positive sequence voltage, frequency, and ROCOF, and in some limited cases, three-phase voltages and currents. We then reflect on the bad data issues that stem from poor recording practices and vague definitions of the PMU status bits to supposedly be used for bad data identification. Finally, the data discovery points to imprecise time stamps with incomplete event start/end time, as well as inconsistent and incomplete event labeling, which combined make the implementation of the data models using supervising learning quite challenging. Following the data discovery study, we hypothesize that because the IC B data has the most complete labels, we should focus our model development on that data and then test it on data from other interconnects. We also define the common metrics used to evaluate the results from the ML algorithm tests. We concluded this section by summarizing the common ML models we used and explaining how we implemented and tested them. The issues from this section are expanded in the Training Dataset Report from this project. The final section of this report deals with the accomplishments and conclusions. As the accomplishments, we formulate the problem we are solving and what is achieved by solving the problem. We then reflect on each of the analytics tools we developed and point out the performance of each tool when applied to solving the mentioned problems. We reference this work for further details to the papers we published on each tool. In the conclusions, we give recommendations on how to improve future PMU recording practices to facilitate the ML algorithm implementation and guidance for the future standardization work aimed at clarifying the ambiguities associated with the PMU status bits. We finally list future tasks that can bring about further improvements in the proposed algorithms. The issues from this section are expanded in the Training, and Test Dataset Report filed at the project completion date.

97 MATHEMATICS AND COMPUTING↗