This package provides functions to estimate the disparities across categories (e.g. Black and white) that persists if a treatment variable (e.g. college) is equalized. Makes estimates by treatment modeling, outcome modeling, and doubly-robust augmented inverse probability weighting estimation, with standard errors calculated by a nonparametric bootstrap. Cross-fitting is supported. Survey weights are supported for point estimation but not for standard error estimation; those applying this package with complex survey samples should consult the data distributor to select an appropriate approach for standard error construction, which may involve calling the functions repeatedly for many sets of replicate weights provided by the data distributor. The methods in this package are described in the accompanying paper: <doi:10.1177/00491241211055769>.
This package provides tools for planning and simulating recurrent event trials with overdispersed count endpoints analyzed using negative binomial (or Poisson) rate models. Implements sample size and power calculations for fixed designs with variable accrual, dropout, maximum follow-up, and event gaps, including methods of Zhu and Lakkis (2014) <doi:10.1002/sim.5947> and Friede and Schmidli (2010) <doi:10.3414/ME09-02-0060> as well as extensions for score-test sizing and gaps between events. Supports group sequential monitoring by building on the gsDesign package. Includes recurrent-event simulation utilities (including seasonal rates), interim data truncation, Wald and score-test inference for rate ratios, and information estimation and sample size re-estimation with or without treatment-group labels.
When added to an existing shiny app, users may subset any developer-chosen R data.frame on the fly. That is, users are empowered to slice & dice data by applying multiple (order specific) filters using the AND (&) operator between each, and getting real-time updates on the number of rows effected/available along the way. Thus, any downstream processes that leverage this data source (like tables, plots, or statistical procedures) will re-render after new filters are applied. The shiny moduleâ s user interface has a minimalist aesthetic so that the focus can be on the data & other visuals. In addition to returning a reactive (filtered) data.frame, IDEAFilter as also returns dplyr filter statements used to actually slice the data.
Estimates adjusted prevalence ratios (PR) and their confidence intervals from logistic regression models, addressing the well-known limitation of odds ratios (OR) as approximations to PR in cross-sectional studies with common outcomes. Supports independent observations (glm()), clustered/multilevel data (glmer() from lme4'), longitudinal data via Generalised Estimating Equations (geeglm() from geepack'), and complex survey designs (svyglm() from survey'). Inference is available via the delta method (conditional and marginal standardisation) and via bootstrap (normal-approximation and percentile intervals). Continuous covariates are handled through user-specified or median-based reference values; flexible baseline specification allows any reference category to be chosen for factor predictors. Based on the methodology described in Amorim & Ospina (2021) <doi:10.1590/0001-3765202120190316>.
Manages, builds and computes statistics and datasets for the construction of quarterly (sub-annual) life tables by exploiting micro-data from either a general or an insured population. References: Pavà a and Lledó (2022) <doi:10.1111/rssa.12769>. Pavà a and Lledó (2023) <doi:10.1017/asb.2023.16>. Pavà a and Lledó (2025) <doi:10.1371/journal.pone.0315937>. Acknowledgements: The authors wish to thank Conselleria de Educación, Universidades y Empleo, Generalitat Valenciana (grants AICO/2021/257; CIAICO/2024/031), Ministerio de Ciencia e Innovación (grant PID2021-128228NB-I00) and Fundación Mapfre (grant Modelización espacial e intra-anual de la mortalidad en España. Una herramienta automática para el calculo de productos de vida') for supporting this research.
Sample surveys use scientific methods to draw inferences about population parameters by observing a representative part of the population, called sample. The SRSWOR (Simple Random Sampling Without Replacement) is one of the most widely used probability sampling designs, wherein every unit has an equal chance of being selected and units are not repeated.This function draws multiple SRSWOR samples from a finite population and estimates the population parameter i.e. total of HT, Ratio, and Regression estimators. Repeated simulations (e.g., 500 times) are used to assess and compare estimators using metrics such as percent relative bias (%RB), percent relative root means square error (%RRMSE).For details on sampling methodology, see, Cochran (1977) "Sampling Techniques" <https://archive.org/details/samplingtechniqu0000coch_t4x6>.
This package implements Bayesian Double-Penalty Tobit Quantile Regression methods for longitudinal interval-censored data as proposed by Zhao et al. (2024) <doi:10.3390/math12121782>. Supports Bayesian Tobit quantile regression with double adaptive Lasso penalty ('PDAL-BTQR'), double Lasso penalty ('PDL-BTQR'), and unpenalized mixed-effects ('P-BTQR'). Handles left, right, interval, and bilateral censoring schemes in longitudinal and clustered structures. Includes Gibbs sampling algorithms, parameter estimation, standard error computation, posterior credible intervals, forecast predictions, DIC, LPML, and diagnostic plotting. References: Tobin (1958) <doi:10.2307/1907382>; Koenker and Bassett (1978) <doi:10.2307/1913643>; Zou (2006) <doi:10.1198/016214506000000735>; Alhamzawi and Yu (2012) <doi:10.1016/j.csda.2011.11.018>; Zhao et al. (2024) <doi:10.3390/math12121782>.
The design of this package allows us to run different clustering packages and compare the results between them, to determine which algorithm behaves best from the data provided. See Martos, L.A.P., Garcà a-Vico, à .M., González, P. et al.(2023) <doi:10.1007/s13748-022-00294-2> "Clustering: an R library to facilitate the analysis and comparison of cluster algorithms.", Martos, L.A.P., Garcà a-Vico, à .M., González, P. et al. "A Multiclustering Evolutionary Hyperrectangle-Based Algorithm" <doi:10.1007/s44196-023-00341-3> and L.A.P., Garcà a-Vico, à .M., González, P. et al. "An Evolutionary Fuzzy System for Multiclustering in Data Streaming" <doi:10.1016/j.procs.2023.12.058>.
This package provides a comprehensive suite of helper functions designed to facilitate the analysis of genomic annotations from the GENCODE database <https://www.gencodegenes.org/>, supporting both human and mouse genomes. This toolkit enables users to extract, filter, and analyze a wide range of annotation features including genes, transcripts, exons, and introns across different GENCODE releases. It provides functionality for cross-version comparisons, allowing researchers to systematically track annotation updates, structural changes, and feature-level differences between releases. In addition, the package can generate high-quality FASTA files containing donor and acceptor splice site motifs, which are formatted for direct input into the MaxEntScan tool (Yeo and Burge, 2004 <doi:10.1089/1066527041410418>), enabling accurate calculation of splice site strength scores.
This package provides an interactive Shiny application and a toolbox of R functions for the management, calculation, filtering, visualization and exploratory analysis of molecular descriptors and ADMET (Absorption, Distribution, Metabolism, Excretion and Toxicity) properties of small molecules. Computes descriptors locally via the Chemistry Development Kit (CDK), and offers drug-likeness filters (Lipinski, Veber, Ghose, Egan, Muegge), the BOILED-Egg model for gastrointestinal absorption and blood-brain barrier permeability, a P-glycoprotein (P-gp, also known as ATP-binding cassette sub-family B member 1, ABCB1) substrate Random Forest classifier, Principal Component Analysis (PCA), t-Distributed Stochastic Neighbor Embedding (t-SNE), Uniform Manifold Approximation and Projection (UMAP), radar plots and Tanimoto / AGglomerative NESting (AGNES) clustering to support compound prioritization in early-stage drug discovery.
This package provides functions to estimate latent dimensions of choice and judgment using Aldrich-McKelvey and Blackbox scaling methods, as described in Poole et al. (2016, <doi:10.18637/jss.v069.i07>). These techniques allow researchers (particularly those analyzing political attitudes, public opinion, and legislative behavior) to recover spatial estimates of political actors ideal points and stimuli from issue scale data, accounting for perceptual bias, multidimensional spaces, and missing data. The package uses singular value decomposition and alternating least squares (ALS) procedures to scale self-placement and perceptual data into a common latent space for the analysis of ideological or evaluative dimensions. Functionality also include tools for assessing model fit, handling complex survey data structures, and reproducing simulated datasets for methodological validation.
This package provides a tool for hydrologic modelling using the Budyko framework and the Dynamic Water Balance model with Dynamical Dimension Search algorithm to calibrate the model and analyze the outputs from interactive graphics. It allows to calculate the water availability in basins and also some water fluxes represented by the structure of the model. See Zhang, L., N., Potter, K., Hickel, Y., Zhang, Q., Shao (2008) <DOI:10.1016/j.jhydrol.2008.07.021> "Water balance modeling over variable time scales based on the Budyko framework - Model development and testing", Journal of Hydrology, 360, 117â 131. See Tolson, B., C., Shoemaker (2007) <DOI:10.1029/2005WR004723> "Dynamically dimensioned search algorithm for computationally efficient watershed model calibration", Water Resources Research, 43, 1â 16.
Takes a distance matrix and plots it as an interactive graph. One point is focused at the center of the graph, around which all other points are plotted in their exact distances as given in the distance matrix. All other non-focus points are plotted as best as possible in relation to one another. Double click on any point to choose a new focus point, and hover over points to see their ID labels. If color label categories are given, hover over colors in the legend to highlight only those points and click on colors to highlight multiple groups. For more information on the rationale and mathematical background, as well as an interactive introduction, see <https://lea-urpa.github.io/focusedMDS.html>.
This is an add-on package to gamlss'. The purpose of this package is to allow users to fit GAMLSS (Generalised Additive Models for Location Scale and Shape) models when the response variable is defined either in the intervals [0,1), (0,1] and [0,1] (inflated at zero and/or one distributions), or in the positive real line including zero (zero-adjusted distributions). The mass points at zero and/or one are treated as extra parameters with the possibility to include a linear predictor for both. The package also allows transformed or truncated distributions from the GAMLSS family to be used for the continuous part of the distribution. Standard methods and GAMLSS diagnostics can be used with the resulting fitted object.
Estimates reference evapotranspiration, crop evapotranspiration, effective rainfall, crop water requirements, root-zone water balance, and irrigation schedules across multiple locations. The calculations use temperature-based procedures described in Food and Agriculture Organization Irrigation and Drainage Paper No. 56 and a workflow inspired by the CROPWAT software for monthly-to-daily interpolation, aggregation into 10-day periods, and irrigation scheduling. Further details of the evapotranspiration calculations are provided by Allen, R.G., Pereira, L.S., Raes, D. and Smith, M. (1998, ISBN:92-5-104219-5) "Crop Evapotranspiration: Guidelines for Computing Crop Water Requirements" <https://www.fao.org/4/X0490E/X0490E00.htm>. The package is an independent implementation and is not affiliated with or endorsed by the Food and Agriculture Organization of the United Nations.
Assesses whether cure models are appropriate for right-censored survival data, where a fraction of subjects may never experience the event of interest. Implements a two-stage workflow combining Kaplan-Meier visualization and comparison of parametric cure and non-cure models by the Akaike information criterion with formal diagnostics for sufficient follow-up and for the presence of a cured fraction. The diagnostics include the statistics of Maller and Zhou (1992) <doi:10.1093/biomet/79.4.731> and Maller and Zhou (1994) <doi:10.1080/01621459.1994.10476889>, the test of Shen (2000) <doi:10.1016/S0167-7152(00)00063-8>, and the ratio estimation of censored uncured subjects ('RECeUS') method of Selukar and Othus (2023) <doi:10.1002/sim.9610>.
This package provides a suite of computer model test functions that can be used to test and evaluate algorithms for Bayesian (also known as sequential) optimization. Some of the functions have known functional forms, however, most are intended to serve as black-box functions where evaluation requires running computer code that reveals little about the functional forms of the objective and/or constraints. The primary goal of the package is to provide users (especially those who do not have access to real computer models) a source of reproducible and shareable examples that can be used for benchmarking algorithms. The package is a living repository, and so more functions will be added over time. For function suggestions, please do contact the author of the package.
Alternative implementation of the beautiful MissForest algorithm used to impute mixed-type data sets by chaining random forests, introduced by Stekhoven, D.J. and Buehlmann, P. (2012) <doi:10.1093/bioinformatics/btr597>. Under the hood, it uses the lightning fast random forest package ranger'. Between the iterative model fitting, we offer the option of using predictive mean matching. This firstly avoids imputation with values not already present in the original data (like a value 0.3334 in 0-1 coded variable). Secondly, predictive mean matching tries to raise the variance in the resulting conditional distributions to a realistic level. This would allow, e.g., to do multiple imputation when repeating the call to missRanger(). Out-of-sample application is supported as well.
This software has evolved from fisheries research conducted at the Pacific Biological Station (PBS) in Nanaimo', British Columbia, Canada. It extends the R language to include two-dimensional plotting features similar to those commonly available in a Geographic Information System (GIS). Embedded C code speeds algorithms from computational geometry, such as finding polygons that contain specified point events or converting between longitude-latitude and Universal Transverse Mercator (UTM) coordinates. Additionally, we include C++ code developed by Angus Johnson for the Clipper library, data for a global shoreline, and other data sets in the public domain. Under the user's R library directory .libPaths()', specifically in ./PBSmapping/doc', a complete user's guide is offered and should be consulted to use package functions effectively.
The package xmapbridge can plot graphs in the X:Map genome browser. X:Map uses the Google Maps API to provide a scrollable view of the genome. It supports a number of species, and can be accessed at http://xmap.picr.man.ac.uk. This package exports plotting files in a suitable format. Graph plotting in R is done using calls to the functions xmap.plot and xmap.points, which have parameters that aim to be similar to those used by the standard plot methods in R. These result in data being written to a set of files (in a specific directory structure) that contain the data to be displayed, as well as some additional meta-data describing each of the graphs.
KnowYourCG (KYCG) is a supervised learning framework designed for the functional analysis of DNA methylation data. Unlike existing tools that focus on genes or genomic intervals, KnowYourCG directly targets CpG dinucleotides, featuring automated supervised screenings of diverse biological and technical influences, including sequence motifs, transcription factor binding, histone modifications, replication timing, cell-type-specific methylation, and trait-epigenome associations. KnowYourCG addresses the challenges of data sparsity in various methylation datasets, including low-pass Nanopore sequencing, single-cell DNA methylomes, 5-hydroxymethylation profiles, spatial DNA methylation maps, and array-based datasets for epigenome-wide association studies and epigenetic clocks (<doi:10.1126/sciadv.adw3027>). KnowYourCG v2, a command-line implementation in C, is available at <https://github.com/zhou-lab/kycg>.
Providing equivalent functions for the dummy classifier and regressor used in Python scikit-learn library. Our goal is to allow R users to easily identify baseline performance for their classification and regression problems. Our baseline models use no predictors, and are useful in cases of class imbalance, multiclass classification, and when users want to quickly identify how much improvement their statistical and machine learning models are over several baseline models. We use a "better" default (proportional guessing) for the dummy classifier than the Python implementation ("prior", which is the most frequent class in the training set). The functions in the package can be used on their own, or introduce methods named dummy_regressor or dummy_classifier that can be used within the caret package pipeline.
It uses the first-order sensitivity index to measure whether the weights assigned by the creator of the composite indicator match the actual importance of the variables. Moreover, the variance inflation factor is used to reduce the set of correlated variables. In the case of a discrepancy between the importance and the assigned weight, the script determines weights that allow adjustment of the weights to the intended impact of variables. If the optimised weights are unable to reflect the desired importance, the highly correlated variables are reduced, taking into account variance inflation factor. The final outcome of the script is the calculated value of the composite indicator based on optimal weights and a reduced set of variables, and the linear ordering of the analysed objects.
This package provides a workflow for correction of Differential Interferometric Synthetic Aperture Radar (DInSAR) atmospheric delay base on Generic Atmospheric Correction Online Service for InSAR (GACOS) data and correction algorithms proposed by Chen Yu. This package calculate the Both Zenith and LOS direction (User Depend). You have to just download GACOS product on your area and preprocessed D-InSAR unwrapped images. Cite those references and this package in your work, when using this framework. References: Yu, C., N. T. Penna, and Z. Li (2017) <doi:10.1016/j.rse.2017.10.038>. Yu, C., Li, Z., & Penna, N. T. (2017) <doi:10.1016/j.rse.2017.10.038>. Yu, C., Penna, N. T., and Li, Z. (2017) <doi:10.1002/2016JD025753>.