This package implements a likelihood-based method for genome polarization, identifying which alleles of SNV markers belong to either side of a barrier to gene flow. The approach co-estimates individual assignment, barrier strength, and divergence between sides, with direct application to studies of hybridization. Includes VCF-to-diem conversion and input checks, support for mixed ploidy and parallelization, and tools for visualization and diagnostic outputs. Based on diagnostic index expectation maximization as described in Baird et al. (2023) <doi:10.1111/2041-210X.14010>.
Distances on dual-weighted directed graphs using priority-queue shortest paths (Padgham (2019) <doi:10.32866/6945>). Weighted directed graphs have weights from A to B which may differ from those from B to A. Dual-weighted directed graphs have two sets of such weights. A canonical example is a street network to be used for routing in which routes are calculated by weighting distances according to the type of way and mode of transport, yet lengths of routes must be calculated from direct distances.
Computes the power and sample size (PASS) required to test for the difference in the mean function between two groups under a repeatedly measured longitudinal or sparse functional design. See the manuscript by Koner and Luo (2023) <https://salilkoner.github.io/assets/PASS_manuscript.pdf> for details of the PASS formula and computational details. The details of the testing procedure for univariate and multivariate response are presented in Wang (2021) <doi:10.1214/21-EJS1802> and Koner and Luo (2023) <arXiv:2302.05612> respectively.
Fits probability models to P- and S-type count data arising from forensic surveys of clothing for the background presence of glass, paint, and related trace material. Built-in models include zeta, zero-inflated zeta, and logarithmic distributions, with a public extension interface for additional models. Inference is available by maximum likelihood, parametric Bayesian methods, the ordinary nonparametric bootstrap, and Rubin's Bayesian Bootstrap. The clothing-survey setting is described by Coulson, Buckleton, Gummer, and Triggs (2001) <doi:10.1016/S1355-0306(01)71847-3>.
This package provides methods and tools designed to improve the forecast accuracy for a linearly constrained multiple time series, while fulfilling the linear/aggregation relationships linking the components (Girolimetto and Di Fonzo, 2024 <doi:10.48550/arXiv.2412.03429>). FoCo2 offers multi-task forecast combination and reconciliation approaches leveraging input from multiple forecasting models or experts and ensuring that the resulting forecasts satisfy specified linear constraints. In addition, linear inequality constraints (e.g., non-negativity of the forecasts) can be imposed, if needed.
This package implements the Genomic Random Interval (GRIN) framework for identifying genomic loci affected by genomic lesions more frequently than expected by chance. Supports multiple lesion classes, lesion constellation analysis, exon-level target-size modeling, genomic lesion visualization, and gene-level association analyses linking genomic lesions or gene expression with binary and time-to-event clinical outcomes. Includes tools for retrieving versioned GRCh38 Ensembl gene, exon, and regulatory-element annotations. The statistical framework is described in Pounds et al. (2013) <doi:10.1093/bioinformatics/btt372>.
Joint mean and dispersion effects models fit the mean and dispersion parameters of a response variable by two separate linear models, the mean and dispersion submodels, simultaneously. It also allows the users to choose either the deviance or the Pearson residuals as the response variable of the dispersion submodel. Furthermore, the package provides the possibility to nest the submodels in one another, if one of the parameters has significant explanatory power on the other. Wu & Li (2016) <doi:10.1016/j.csda.2016.04.015>.
This package provides an interface to the Mapbox GL JS (<https://docs.mapbox.com/mapbox-gl-js/guides>) and the MapLibre GL JS (<https://maplibre.org/maplibre-gl-js/docs/>) interactive mapping libraries to help users create custom interactive maps in R. Users can create interactive globe visualizations; layer sf objects to create filled maps, circle maps, heatmaps', and three-dimensional graphics; and customize map styles and views. The package also includes utilities to use Mapbox and MapLibre maps in Shiny web applications.
This package implements methods for estimating generalized estimating equations (GEE) with advanced options for flexible modeling and handling missing data. This package provides tools to fit and analyze GEE models for longitudinal data, allowing users to address missingness using a variety of imputation techniques. It supports both univariate and multivariate modeling, visualization of missing data patterns, and facilitates the transformation of data for efficient statistical analysis. Designed for researchers working with complex datasets, it ensures robust estimation and inference in longitudinal and clustered data settings.
Calculate a multivariate functional principal component analysis for data observed on different dimensional domains. The estimation algorithm relies on univariate basis expansions for each element of the multivariate functional data (Happ & Greven, 2018) <doi:10.1080/01621459.2016.1273115>. Multivariate and univariate functional data objects are represented by S4 classes for this type of data implemented in the package funData'. For more details on the general concepts of both packages and a case study, see Happ-Kurz (2020) <doi:10.18637/jss.v093.i05>.
Asymptotic efficient closed-form estimators (MLEces) are provided in this package for three multivariate distributions(gamma, Weibull and Dirichlet) whose maximum likelihood estimators (MLEs) are not in closed forms. Closed-form estimators are strong consistent, and have the similar asymptotic normal distribution like MLEs. But the calculation of MLEces are much faster than the corresponding MLEs. Further details and explanations of MLEces can be found in. Jang, et al. (2023) <doi:10.1111/stan.12299>. Kim, et al. (2023) <doi:10.1080/03610926.2023.2179880>.
This ONEST software implements the method of assessing the pathologist agreement in reading PD-L1 assays (Reisenbichler et al. (2020 <doi:10.1038/s41379-020-0544-x>)), to determine the minimum number of evaluators needed to estimate agreement involving a large number of raters. Input to the program should be binary(1/0) pathology data, where â 0â may stand for negative and â 1â for positive. Additional examples were given using the data from Rimm et al. (2017 <doi:10.1001/jamaoncol.2017.0013>).
SPINA (Structure Parameter Inference Approach) is a methodology to calculate constant structure parameters of endocrine homeostatic systems from steady-state hormone and metabolite concentrations. Methods and equations for thyroid homeostasis (SPINA Thyr) have been described in Dietrich et al. (2012) <doi:10.1155/2012/351864> and Dietrich et al. (2016) <doi:10.3389/fendo.2016.00057>, and for glucose homeostasis (SPINA Carb) in Dietrich et al. (2022) <doi:10.1038/s41598-022-22531-3> and Dietrich et al. (2024) <doi:10.1111/1753-0407.13525>.
Augmenting a matched data set by generating multiple stochastic, matched samples from the data using a multi-dimensional histogram constructed from dropping the input matched data into a multi-dimensional grid built on the full data set. The resulting stochastic, matched sets will likely provide a collectively higher coverage of the full data set compared to the single matched set. Each stochastic match is without duplication, thus allowing downstream validation techniques such as cross-validation to be applied to each set without concern for overfitting.
This software is useful for loading .fasta or .gbk files, and for retrieving sequences from GenBank dataset <https://www.ncbi.nlm.nih.gov/genbank/>. This package allows to detect differences or asymmetries based on nucleotide composition by using local linear kernel smoothers. Also, it is possible to draw inference about critical points (i. e. maximum or minimum points) related with the derivative curves. Additionally, bootstrap methods have been used for estimating confidence intervals and speed computational techniques (binning techniques) have been implemented in seq2R'.
Inferring causation from spatial cross-sectional data through empirical dynamic modeling (EDM), with methodological extensions including geographical convergent cross mapping from Gao et al. (2023) <doi:10.1038/s41467-023-41619-6>, geographical cross mapping cardinality as introduced by Lyu et al. (2026) <doi:10.1080/13658816.2026.2687121>, as well as the spatial causality test following the approach of Herrera et al. (2016) <doi:10.1111/pirs.12144>, together with geographical pattern causality proposed in Zhang & Wang (2025) <doi:10.1080/13658816.2025.2581207>.
Estimates sampling errors and produces indicator tables for complex survey data. Supports weighted totals, proportions, standard errors, confidence intervals, coefficients of variation, design effects, unweighted frequencies, grouped estimates, domain estimates, optional stratification and clustering variables, and customizable exports to .xlsx files. Survey estimation is based on design-based inference using Taylor series linearization implemented in the survey package (Lumley, 2004, <doi:10.18637/jss.v009.i08>; Lumley, 2010, ISBN:9780470284308). The package provides a reproducible workflow for official statistics, household surveys, and applied survey research.
This package provides helper functions and wrappers to simplify authentication, data retrieval, and result processing from the VALD APIs'. Designed to streamline integration for analysts and researchers working with VALD's external APIs'. For further documentation on integrating with VALD APIs', see: <https://support.vald.com/hc/en-au/articles/23415335574553-How-to-integrate-with-VALD-APIs>. For a step-by-step guide to using this package, see: <https://support.vald.com/hc/en-au/articles/48730811824281-A-guide-to-using-the-valdr-R-package>.
This package provides tools to convert statistical analysis objects from R into tidy data frames, so that they can more easily be combined, reshaped and otherwise processed with tools like dplyr, tidyr and ggplot2. The package provides three S3 generics: tidy, which summarizes a model's statistical findings such as coefficients of a regression; augment, which adds columns to the original data such as predictions, residuals and cluster assignments; and glance, which provides a one-row summary of model-level statistics.
Comprehensive set of tools for performing system identification of both linear and nonlinear dynamical systems directly from data. The Automatic Regression for Governing Equations (ARGOS) simplifies the complex task of constructing mathematical models of dynamical systems from observed input and output data, supporting various types of systems, including those described by ordinary differential equations. It employs optimal numerical derivatives for enhanced accuracy and employs formal variable selection techniques to help identify the most relevant variables, thereby enabling the development of predictive models for system behavior analysis.
This package provides a framework for the identification and analysis of Differentially Expressed Single Nucleotide Polymorphisms (deSNPs) using high-throughput sequencing data. It enables users to import SNP count data from variant files, perform allele-specific read count extraction, and statistically detect SNPs showing significant differences in allele expression between biological conditions or sample groups. This package contains tools for calculating SNP-index and Delta SNP-index from VCF-derived allele depth data with statistical testing and filtering, including sliding-window analysis of genomic regions.
Analyses gene expression data derived from microarray experiments to detect differentially expressed genes (DEGs) by employing majority voting across five statistical models: Welch t-test, one-way ANOVA, Dunnett's test, Half's modified t-test, and the Wilcoxon-Mann-Whitney U-test. Combined p-values are computed with Fisher's method. Gene annotation is optional: users may supply a GEO SOFT annotation table or rely on row names directly. Boyer, R.S., Moore, J.S. (1991) <doi:10.1007/978-94-011-3488-0_5>.
This package provides coefficients of interrater reliability that are generalized to cope with randomly incomplete (i.e. unbalanced) datasets without any imputation of missing values or any (row-wise or column-wise) omissions of actually available data. Applied to complete (balanced) datasets, these generalizations yield the same results as the common procedures, namely the Intraclass Correlation according to McGraw & Wong (1996) \doi10.1037/1082-989X.1.1.30 and the Coefficient of Concordance according to Kendall & Babington Smith (1939) \doi10.1214/aoms/1177732186.
Latent binary Bayesian neural networks (LBBNNs) are implemented using torch', an R interface to the LibTorch backend. Supports mean-field variational inference as well as flexible variational posteriors using normalizing flows. The standard LBBNN implementation follows Hubin and Storvik (2024) <doi:10.3390/math12060788>, using the local reparametrization trick as in Skaaret-Lund et al. (2024) <https://openreview.net/pdf?id=d6kqUKzG3V>. Input-skip connections are also supported, as described in Høyheim et al. (2025) <doi:10.48550/arXiv.2503.10496>.