SEARCH · Engineering Papers
Results for “152 BASIC BIOLOGICAL SCIENCES”
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Graph-based machine learning improves just-in-time defect prediction
The increasing complexity of today’s software requires the contribution of thousands of developers. This complex collaboration structure makes developers more likely to introduce defect-prone changes that lead to software faults. Determining when these defect-prone changes are introduced has proven challenging, and using traditional machine learning (ML) methods to make these determinations seems to have reached a plateau. In this work, we build contribution graphs consisting of developers and source files to capture the nuanced complexity of changes required to build software. By leveraging these contribution graphs, our research shows the potential of using graph-based ML to improve Just-In-Time (JIT) defect prediction. We hypothesize that features extracted from the contribution graphs may be better predictors of defect-prone changes than intrinsic features derived from software characteristics. We corroborate our hypothesis using graph-based ML for classifying edges that represent defect-prone changes. This new framing of the JIT defect prediction problem leads to remarkably better results. We test our approach on 14 open-source projects and show that our best model can predict whether or not a code change will lead to a defect with an F1 score as high as 77.55% and a Matthews correlation coefficient (MCC) as high as 53.16%. This represents a 152% higher F1 score and a 3% higher MCC over the state-of-the-art JIT defect prediction. We describe limitations, open challenges, and how this method can be used for operational JIT defect prediction.
Publication and Impact of Preprints Included in the First 100 Editions of the CDC COVID-19 Science Update: Content Analysis
Preprints are publicly available manuscripts posted to various servers that have not been peer reviewed. Although preprints have existed since 1961, they have gained increased popularity during the COVID-19 pandemic due to the need for immediate, relevant information. The aim of this study is to evaluate the publication rate and impact of preprints included in the Centers for Disease Control and Prevention (CDC) COVID-19 Science Update and assess the performance of the COVID-19 Science Update team in selecting impactful preprints. All preprints in the first 100 editions (April 1, 2020, to July 30, 2021) of the Science Update were included in the study. Preprints that were not published were categorized as “unpublished preprints.” Preprints that were subsequently published exist in 2 versions (in a peer-reviewed journal and on the original preprint server), which were analyzed separately and referred to as “peer-reviewed preprint” and “original preprint,” respectively. Time to publish was the time interval between the date on which a preprint was first posted and the date on which it was first available as a peer-reviewed article. Impact was quantified by Altmetric Attention Score and citation count for all available manuscripts on August 6, 2021. Preprints were analyzed by publication status, publication rate, preprint server, and time to publication. Of the 275 preprints included in the CDC COVID-19 Science Update during the study period, most came from three servers: medRxiv (n=201, 73.1%), bioRxiv (n=41, 14.9%), and SSRN (n=25, 9.1%), with 8 (2.9%) coming from other sources. Additionally, 152 (55.3%) were eventually published. The median time to publish was 2.3 (IQR 1.4-3.7). When preprints posted in the last 2.3 months were excluded (to account for the time to publish), the publication rate was 67.8%. Moreover, 76 journals published at least one preprint from the CDC COVID-19 Science Update, and 18 journals published at least three. The median Altmetric Attention Score for unpublished preprints (n=123, 44.7%) was 146 (IQR 22-552) with a median citation count of 2 (IQR 0-8); for original preprints (n=152, 55.2%), these values were 212 (IQR 22-1164) and 14 (IQR 2-40), respectively; for peer-review preprints, these values were 265 (IQR 29-1896) and 19 (IQR 3-101), respectively. Prior studies of COVID-19 preprints found publication rates between 5.4% and 21.1%. Preprints included in the CDC COVID-19 Science Update were published at a higher rate than overall COVID-19 preprints, and those that were ultimately published were published within months and received higher attention scores than unpublished preprints. These findings indicate that the Science Update process for selecting preprints had a high fidelity in terms of their likelihood to be published and their impact. The incorporation of high-quality preprints into the CDC COVID-19 Science Update improves this activity’s capacity to inform meaningful public health decision-making.
Draft genome of multiple resistance donor plant Sinapis alba: An insight into SSRs, annotations and phylogenetics
Sinapis alba is a wild member of the Brassicaceae family reported to possess genetic resistance against major biotic and abiotic stresses of oilseed brassicas. However, the resistance nature of S. alba was not exploited generously due to the unavailability of usable genome sequences in public databases. Therefore, the present study was conducted to assemble the first draft genome from raw whole genome shotgun sequences with annotation and develop simple sequence repeat markers for molecular genetics and marker-assisted breeding. Results The raw genome sequences had 96x coverage on the Illumina platform with 170 Gbp data. The developed assembly by SOAPdenovo2 has ~459 Mbp genome size covered in 403,423 contigs with an average size of 1138.04 bp. The assembly was BLASTX with Arabidopsis thaliana which showed 32.9% positive hits between both plants. The top hit species distribution analysis showed the highest similarity with A. thaliana. A total of 809,597 GO level annotations were recorded after BLASTX results, and 34,012 sequences were annotated with different enzyme codes grouped under seven classes. The gene prediction tool AUGUSTUS identified 113,107 probable genes with an average size of 684 bp. The biochemical pathway annotation assigned 16,119 potential genes to 152 KEGG maps and 1751 enzyme codes. The development of potential SSRs from the de-novo assembly yielded 70731 unique primer pairs. Out of 159 randomly selected SSR markers for validation, 149 successfully amplified in S. alba. However, 10 SSR markers did not amplify during the validation experiment. Conclusion The annotated genome assembly with a large number of SSRs was developed in the present study. To the best of our knowledge, this is the first report of S. alba genome assembly development, annotation, and SSRs mining to date. The data presented here will be a very important resource for future crop improvement programs, especially for resistant breeding.
Ion Mobility Separations Using Cocentric Architecture
Ion mobility separations are usually performed in linear channels, which, when extended, can have a large footprint. In this work, we explored the performance of an ion mobility device with a curved architecture which can have a more compact form. The Co-centric Ion Mobility Spectrometer (CIMS) works by manipulating ions between two co-centric surfaces, each containing a serpentine track. The mobility separation inside CIMS is achieved using traveling waveforms (TWs). We initially evaluated the device using ion trajectory simulations using SIMION, which indicated that when ions traveled circularly inside CIMS, they resulted in similar resolving powers and transmitted m/z range as traveling in a straight path in structures for lossless ion manipulations (SLIM). We then performed experimental validation of CIMS in conjunction with a TOF MS. The CIMS was made of 2 flexible printed circuit board materials folded into concentric cylinders separated by a gap of 2.8 mm. The device was about 50 mm diameter × 152 mm long and provided 1.846 m of serpentine path length. Three sets of mixtures (Agilent tune mixture, tetraalkylammonium salts, and 8 peptide mixture) and four traveling waveform profiles (square, sine, triangle, and sawtooth) were used. The sawtooth TW profile produced a slightly higher resolving power for the Agilent tuning mixture and tetraalkylammonium ions. The average resolving power for Agilent tune mixture ions ranged from 37 (using sawtooth TW) to 27 (using square TW). For tetraalkylammonium ions, the average resolving powers ranged from 45 (sawtooth TW) to 31 (square TW). For the peptide mixture ions, the resolving power was similar among the four TW profiles and ranged from 51 to 56. The average percent error in TW CCS for the peptide mixture ions ranged was about 0.4%. In conclusion, the new device showed promising results for a device made of a flexible printed circuit board material, but improvements are needed to further increase the resolving power.