Related papers: Non-destructive methods for assessing tree fiber l…
In order for biomass drying processes to be efficient, it is crucial to achieve the target residual water content within a close margin, since more conservative drying would result in a waste of energy. A method for a reliable estimation of…
Selective inference is considered for testing trees and edges in phylogenetic tree selection from molecular sequences. This improves the previously proposed approximately unbiased test by adjusting the selection bias when testing many trees…
Regression trees and their ensemble methods are popular methods for nonparametric regression: they combine strong predictive performance with interpretable estimators. To improve their utility for locally smooth response surfaces, we study…
The use of optical fiber as sensor as well as transmission medium for sensing data is discussed, enabling the use of optically active sensors without power supply at distances of tens of kilometers. Depending on the interrogation system, a…
It is widely assumed that $O(m+\lg \sigma)$ is the best one can do for finding a pattern of length $m$ in a compacted trie storing strings over an alphabet of size $\sigma$, if one insists on linear-size data structures and deterministic…
Data augmentation is widely used for training a neural network given little labeled data. A common practice of augmentation training is applying a composition of multiple transformations sequentially to the data. Existing augmentation…
We investigate lensless endoscopy using coherent beam combining and aperiodic multicore fibers (MCF). We show that diffracted orders, inherent to MCF with periodically arranged cores, dramatically reduce the field of view (FoV) and that…
In recent years, dynamically growing data and incrementally growing number of classes pose new challenges to large-scale data classification research. Most traditional methods struggle to balance the precision and computational burden when…
Fiber tractography on diffusion imaging data offers rich potential for describing white matter pathways in the human brain, but characterizing the spatial organization in these large and complex data sets remains a challenge. We show that…
Many existing interpretation methods are based on Partial Dependence (PD) functions that, for a pre-trained machine learning model, capture how a subset of the features affects the predictions by averaging over the remaining features.…
A correlation optical time-domain reflectometry (COTDR) method is presented, which measures the propagation delay with an accuracy of a few picoseconds. This accuracy is achieved using a test signal data rate of 10 Gbit/s and employing…
Decision trees are one of the most popular classifiers in the machine learning literature. While the most common decision tree learning algorithms treat data as a batch, numerous algorithms have been proposed to construct decision trees…
Anatomic tracing data provides detailed information on brain circuitry essential for addressing some of the common errors in diffusion MRI tractography. However, automated detection of fiber bundles on tracing data is challenging due to…
In this paper, we present a distributed algorithm to compute various parameters of a tree such as the process number, the edge search number or the node search number and so the pathwidth. This algorithm requires n steps, an overall…
Depth First Search (DFS) tree is a fundamental data structure for solving graph problems. The DFS tree of a graph $G$ with $n$ vertices and $m$ edges can be built in $O(m+n)$ time. Till date, only a few algorithms have been designed for…
In this paper we propose a simple method to reject the high-frequency noise in the evaluation of statistical uncertainty of coherent optical fiber links. Specifically, we propose a preliminary data filtering, separated from the frequency…
Fiber photometry permits monitoring fluorescent indicators of neural activity in behaving animals. Optical fibers are typically used to excite and collect fluorescence from genetically-encoded calcium indicators expressed by a subset of…
Label distribution learning (LDL) is a general learning framework, which assigns to an instance a distribution over a set of labels rather than a single label or multiple labels. Current LDL methods have either restricted assumptions on the…
A crucial problem in genome assembly is the discovery and correction of misassembly errors in draft genomes. We develop a method that will enhance the quality of draft genomes by identifying and removing misassembly errors using paired…
This paper proposes a novel approach for statistical modelling of a continuous random variable $X$ on $[0, 1)$, based on its digit representation $X=.X_1X_2\ldots$. In general, $X$ can be coupled with a latent random variable $N$ so that…