| Software | Resource | Containers | Description | AI Tags | AI Research Discipline | AI Software Type | Documentation, Uses, and more |
|---|---|---|---|---|---|---|---|
| abaqus | mcc | N/A | SIMULIA Abaqus 2026 (Established Products) is a finite element analysis suite for advanced structural and multiphysics simulation. It includes Abaqus/Standard (implicit) and Abaqus/Explicit solvers, the Abaqus/CAE modeling environment, ODB API, and cosimulation services, handling nonlinear static, dynamic, thermal, contact, and coupled analyses. It provides the abaqus and abq2026 commands. Used across mechanical, civil, automotive, and aerospace engineering research. | Simulation, Finite Element Analysis, Computer-Aided Engineering, Materials Analysis | Mechanical Engineering | Simulation Software | Documentation, Uses, and more |
| abricate | lcc | 1.0 | ABRicate, version 1.0.1, performs mass in-silico screening of assembled bacterial genome contigs for acquired antimicrobial-resistance (AMR) and virulence genes. It runs BLAST-based nucleotide searches of the input assembly against a chosen bundled database — including NCBI AMRFinderPlus, CARD, ResFinder, ARG-ANNOT, MEGARES, EcOH, PlasmidFinder, and VFDB — reporting each hit's gene, percent coverage, percent identity, contig location, and database accession. It only accepts assembled sequence (FASTA of contigs), not raw reads, and does not detect point-mutation-mediated resistance. In addition to per-genome screening against a selected database, it can list the installed databases and collapse many per-genome reports into a presence/absence summary matrix across a collection of isolates for comparative surveillance. | Antimicrobial Resistance, Genomic Analysis, Bacterial Genomics | Biology, Biological Sciences | Tool | Documentation, Uses, and more |
| absence-aware-clustering | mcc | 1.0 | Absence-Aware-Clustering (AAC), from the Raphael group at Princeton, is a Snakemake-based preprocessing tool for tumor phylogenetics that clusters somatic mutations from multi-sample (e.g. longitudinal or multi-region) cancer sequencing data while explicitly modeling the presence or absence of each mutation in each sample. Standard mutation-clustering approaches can be confounded when a mutation is truly absent (copy number zero or not present in a sample) versus merely low frequency; AAC accounts for this so that the resulting mutation clusters better reflect the underlying clonal populations. It wraps PyClone-VI (a variational-inference clustering of variant allele frequencies with copy-number correction) inside a Snakemake workflow, and its clustered output serves as input to downstream tumor phylogeny reconstruction. In this deployment it runs in a conda environment named 'aac' containing snakemake and pyclone-vi, and it is packaged alongside CALDER for reconstructing tumor phylogenies from longitudinal sequencing on the MCC cluster. | Documentation, Uses, and more | |||
| abyss | lcc | N/A | ABySS is a de novo, parallel, paired-end sequence assembler that is designed for short reads. The single-processor version is useful for assembling genomes up to 100 Mbases in size. The parallel version is implemented using MPI and is capable of assembling larger genomes. Description Source: https://www.bcgsc.ca/resources/software/abyss |
Assembly, De Novo Assembly, Sequence Assembler, Large Genomes | Biological Sciences | Hpc Tool | Documentation, Uses, and more |
| adapterremoval | lcc | 1.0 | AdapterRemoval is a read-preprocessing tool that removes residual adapter sequences and trims low-quality bases from high-throughput (Illumina) sequencing reads, and it can identify and merge overlapping paired-end reads into single consensus fragments. It handles both single- and paired-end data, can trim ambiguous (N) bases and low-quality termini, discard reads that become too short, and reconstruct/collapse read pairs — a feature especially valued in ancient-DNA and short-insert library work where read merging improves downstream analysis. It can also infer adapter sequences directly from the data when they are unknown. Inputs are FASTQ files (gzip supported); outputs are trimmed/collapsed FASTQ files plus a settings/summary report. This is version 2.3.2 from bioconda. | Bioinformatics, Hts Data Processing, Sequence Analysis | Biology, Biological Sciences | Bioinformatics Tool | Documentation, Uses, and more |
| adios | lcc | N/A | The Adaptable IO System (ADIOS) | I/O, High Performance Computing, Data Management | High Performance Computing, Scientific Data Management | Library | Documentation, Uses, and more |
| adios2 | mcc, lcc | N/A | ADIOS2 is developed as part of the United States Department of Energy's Exascale Computing Project. It is a framework for scientific data I/O to publish and subscribe to data when and where required. ADIOS2 transports data as groups of self-describing variables and attributes across different media types (such as files, wide-area-networks, and remote direct memory access) using a common application programming interface for all transport modes. ADIOS2 can be used on supercomputers, cloud systems, and personal computers. | I/O Middleware, High-Performance Computing, Data Management, Parallel Computing | Computer Science, Computer & Information Sciences | Library | Documentation, Uses, and more |
| admixture | mcc, lcc | 1.0 | ADMIXTURE (bioconda 1.3.0) is a fast population-genetics program that estimates individual ancestry proportions from multilocus SNP genotype data under the same statistical model as STRUCTURE, but using an efficient block-relaxation numerical optimization for maximum-likelihood estimation, making it far faster on genome-scale datasets. Given a user-specified number of ancestral populations K, it produces a Q matrix of per-individual ancestry fractions and a P matrix of population allele frequencies. It supports cross-validation (the --cv flag) to help choose the best K by comparing prediction error across values, and offers a supervised mode and bootstrap standard errors. Input is typically PLINK format (.bed with accompanying .bim/.fam, or .ped), and analyses commonly iterate over several K values, selecting the one that minimizes cross-validation error. The Q output is commonly visualized as stacked bar (structure) plots and can be post-processed with tools like CLUMPP and pong. | genetics, population genetics, bioinformatics | Biological Sciences | Statistical Software | Documentation, Uses, and more |
| advisor | mcc, lcc, ecc | N/A | Intel Advisor is a performance analysis tool designed to help developers optimize and parallelize their software on Intel architectures. It provides insights for efficient vectorization, threading, and offloading to accelerators, aiding in achieving maximum performance from applications. | Hpc, Performance Optimization, Parallel Computing | Engineering & Technology | Performance Optimization | Documentation, Uses, and more |
| aflow | mcc, lcc | 1.0 | AFLOW (Automatic FLOW for Materials Discovery, built from source at tag v4.0.5 with the aflow_prototype_encyclopedia submodule) is an automated framework for high-throughput first-principles computational materials science. It generates and manages large batches of DFT calculations (primarily orchestrating VASP), performs crystal symmetry analysis and standardization, and computes structural, thermodynamic, elastic, electronic, and thermal properties that feed the AFLOW.org materials database. It ships an extensive library of crystal-structure prototypes (the prototype encyclopedia) that can be enumerated and decorated with chemistries to seed screening campaigns, and it implements methods such as AGL/AEL (Debye-model thermal/elastic) and the AFLOW-CHULL convex-hull thermodynamic stability analysis. It is driven through the aflow command, which can build prototypes, run symmetry analysis, and drive workflow generation. This build compiles against a from-source libcurl 8.10.1 on Rocky 9 to support its network-enabled database features. | Computational Software, Materials Science, Density Functional Theory | Physics & Physical Sciences, Physical Sciences | Materials Modeling & Simulation | Documentation, Uses, and more |
| afterqc | lcc | 1.0 | AfterQC is an automatic quality-control, filtering, trimming, and error-correction tool for FASTQ data from Illumina HiSeq/MiSeq-style sequencing runs, aimed at requiring essentially no configuration. It automatically detects and trims sequencing adapters and bubble/artifact regions, discards low-quality and polluted reads, and — uniquely among basic QC tools — performs overlap-based error correction on paired-end reads by using the overlap of read pairs to fix mismatched bases. It profiles the data before and after filtering and emits per-file HTML reports with quality, base-content, and k-mer plots plus JSON summaries, writing cleaned FASTQ files into good/bad/QC output folders, with options to control quality thresholds and trimming. It is a predecessor to the same author's later fastp. This is bioconda afterqc version 0.9.7. | Quality Control, High-Throughput Sequencing Data, Next-Generation Sequencing | Biology, Biological Sciences | Application | Documentation, Uses, and more |
| agat | lcc | 1.0 | AGAT (Another GTF/GFF Analysis Toolkit), version 1.0.0, is a comprehensive suite of Perl scripts for parsing, validating, standardizing, and manipulating genome-annotation files in GTF/GFF (all flavors and versions). Its parser rebuilds a complete, consistent feature hierarchy (gene → mRNA → exon/CDS/UTR), fixing common problems such as missing parent features, duplicate or out-of-order records, and inconsistent identifiers, so downstream tools receive clean input. The toolkit exposes dozens of agat_sp_* and agat_sq_* scripts for tasks like extracting sequences, converting between formats, filtering by feature type or attribute, merging annotations, computing summary statistics, and complementing/adding UTRs or introns. A common first step is to normalize an annotation by converting it through the toolkit's GxF-to-GxF cleaner before further processing. It is a de-facto standard for wrangling gene-model files in eukaryotic annotation pipelines. | Genome Assembly, Annotation, Evaluation, Genomic Analysis | Genomics, Biological Sciences | Pipeline | Documentation, Uses, and more |
| agptools | mcc | 1.0 | agptools (installed from the esrice/agptools GitHub repository via pip) is a lightweight command-line toolkit for programmatically editing AGP (A Golden Path) files, the standard tab-delimited format that specifies how component sequences (contigs) are ordered, oriented, and gapped to build scaffolds or chromosomes. It provides subcommands for common scaffold-manipulation operations during genome-assembly curation — for example splitting a scaffold at a given point, joining scaffolds together, flipping (reverse-complementing) a component or scaffold's orientation, and removing or renaming components — so that curators can adjust an assembly's structure by editing the AGP rather than rewriting FASTA. This is especially useful when incorporating manual corrections (e.g., from Hi-C contact-map review) into an assembly in a documented, reproducible way. Inputs are an AGP file (and, for producing sequence, the corresponding component FASTA); outputs are the modified AGP that can then be rendered into an updated assembly FASTA. It fits into scaffolding/curation pipelines alongside tools like RagTag and SALSA2 where AGP is the interchange format for assembly structure. | Genome Assembly | Bioinformatics | Documentation, Uses, and more | |
| ai-chat-openwebui | lcc, ecc | 1.0 | AI-Chat-OpenWebUI is an Open OnDemand interactive-app registration that launches Open WebUI, a self-hosted, browser-based chat interface for interacting with large language models. Open WebUI provides a ChatGPT-style experience — conversation history, model selection, prompt management, and multi-user access — running against locally hosted models (e.g., via Ollama or OpenAI-compatible backends) so that data stays within the cluster. This deployment adds retrieval-augmented generation (RAG) support, letting users upload or reference document collections that the model consults to ground its answers in local knowledge. In the SDS catalog this entry is a stub registering the ECC cluster's Open OnDemand browser-based interactive app rather than a built container image; users launch it as an OOD session and connect through their browser. | Documentation, Uses, and more | |||
| allmaps | lcc | N/A | ALLMAPS is the genome-scaffolding tool from the JCVI utility library, which integrates multiple genetic and physical maps to order and orient assembly scaffolds into chromosomes. On LCC it is a conda environment (/opt/ohpc/pub/libs/conda/env/allmaps) loaded via ccs/allmaps that bundles jcvi 1.0.6 along with its dependencies: the Concorde TSP solver (the 'concorde' binary ALLMAPS uses for scaffold ordering), UCSC-style helpers faSize and liftOver, faidx, gffutils-cli, and ete3. It is invoked through the jcvi assembly.allmaps module and is used in genome-assembly and comparative-genomics workflows. | Genomics, chromosome reconstruction, bioinformatics | Genome Assembly, Bioinformatics | assembly scaffolding tool | Documentation, Uses, and more |
| alphafold3 | lcc | 1.0 | AlphaFold 3 is DeepMind's deep-learning system for predicting the joint three-dimensional structure of biomolecular complexes. Unlike AlphaFold 2, which focused on single-chain and multimeric proteins, AlphaFold 3 predicts structures for a mixture of molecule types together — proteins, DNA, RNA, small-molecule ligands, ions, and covalent/post-translational modifications — using a diffusion-based generative architecture over atomic coordinates. Input is a JSON file specifying the sequences and entities of the complex; the pipeline first runs a data/feature stage that generates multiple sequence alignments (using the bundled HMMER 3.4, e.g. jackhmmer against sequence databases) and searches for structural templates, then runs GPU-accelerated inference in JAX/CUDA to produce ranked predicted structures in mmCIF format together with confidence metrics (pLDDT, PAE, and ranking scores). Running it requires the large genetic-sequence databases and, importantly, obtaining the model weights directly from DeepMind under their usage terms. This deployment is v3.0.1 built from the tagged upstream release. | Protein-Structure, Deep-Learning, Bioinformatics | Bioinformatics, Biochemistry and Molecular Biology | Predictive Modeling | Documentation, Uses, and more |
| amas | mcc | 1.0 | AMAS (Alignment Manipulation And Summary) is a fast, lightweight tool for computing summary statistics on and manipulating multiple sequence alignments, widely used in phylogenomics to characterize and assemble large multi-locus datasets. Its summary mode reports per-alignment and per-taxon metrics such as alignment length, number of taxa, proportion of missing/gap characters, number and proportion of parsimony-informative sites, GC content, and variable sites. Beyond QC it can concatenate many single-locus alignments into a supermatrix (writing partition files), split alignments, convert between formats (FASTA, PHYLIP, NEXUS), and remove taxa or sites. It runs as a Python command line exposing these summary, concat, and related subcommands. This is version 1.0. | Documentation, Uses, and more | |||
| amber | mcc, lcc | 1.0 | Amber 20 (Assisted Model Building with Energy Refinement) - a suite of biomolecular molecular-dynamics programs (pmemd, sander, tleap, cpptraj, antechamber) for simulating proteins, nucleic acids, carbohydrates and small molecules: MD, energy minimization, free-energy and trajectory analysis. This build: GNU toolchain + OpenMPI/OpenMP (CPU). | Molecular Dynamics, Biomolecular Simulations, Quantum Chemistry | Biochemistry & Molecular Biology, Biological Sciences | Molecular Dynamics | Documentation, Uses, and more |
| ambertools | mcc, lcc | 1.0 | AmberTools 21 - the free, open component of Amber: tleap and antechamber (system and parameter prep), sander (MD and minimization), cpptraj and pytraj (trajectory analysis), MMPBSA.py (free-energy), plus force-field and QM/MM utilities for biomolecular simulation. Provided as a conda environment. | Molecular Dynamics, Computational Chemistry, Biomolecular Structure, Energy Minimization | Biophysics, Biological Sciences | Simulation Software | Documentation, Uses, and more |
| amdblis | mcc | N/A | AMD Optimized BLIS is a portable software framework for instantiating high-performance BLAS-like dense linear algebra libraries. | Computational Software, Data Analysis, Visualization | Biophysics, Biological Sciences | Computational Tool | Documentation, Uses, and more |
| amdfftw | mcc | N/A | An AMD optimized version of FFTW. FFTW is a C subroutine library for computing the discrete Fourier transform (DFT) in one or more dimensions, of arbitrary input size, and of both real and complex data (as well as of even/odd data, i.e. the discrete cosine/sine transforms or DCT/DST). We believe that FFTW, which is free software, should become the FFT library of choice for most applications. Description Source: https://www.fftw.org/ |
Mathematics | Library | Documentation, Uses, and more | |
| amdlibflame | mcc | N/A | libFLAME AMD Optimized version is a portable library for dense matrix computations, providing much of the functionality present in Linear Algebra Package LAPACK. It includes a compatibility layer, FLAPACK, which includes complete LAPACK implementation. | Linear Algebra, Numerical Computations, Performance Optimization | Mathematics, Computer & Information Sciences | Library | Documentation, Uses, and more |
| amdlibm | mcc | N/A | AMD LibM is a software library containing a collection of basic math functions optimized for x86-64 processor-based machines. It provides many routines from the list of standard C99 math functions. Applications can link into AMD LibM library and invoke math functions instead of compilers math functions for better accuracy and performance. | Mathematics, Numerical Computation, Math Library | Mathematics, Engineering & Technology | Math Library | Documentation, Uses, and more |
| amdscalapack | mcc | N/A | ScaLAPACK is a library of high-performance linear algebra routines for parallel distributed memory machines. It depends on external libraries including BLAS and LAPACK for Linear Algebra computations. | Mathematics | Library | Documentation, Uses, and more | |
| anaconda | lcc | N/A | Anaconda3 is a popular open-source distribution of Python and R for scientific computing, data science, and machine learning. It simplifies package management and deployment by bundling numerous libraries and tools, and it includes the conda package manager, which makes installing, running, and updating various packages and their dependencies convenient. | Data Science, Scientific Computing, Package Management, Python, R | Computer and Information Sciences | Package Management | Documentation, Uses, and more |
| ancestry_hmm | lcc | 1.0 | Ancestry_HMM infers local ancestry along the chromosomes of admixed individuals, that is, it assigns each genomic segment to the ancestral population it descends from, using a hidden Markov model whose hidden states are local ancestry configurations and whose transitions depend on the genetic map and the time since admixture. Rather than requiring hard genotype calls, it can take read counts and allele frequencies of the source populations as input, making it suitable for low-coverage data, and it jointly models the admixture proportions and timing to produce posterior probabilities of each ancestry state at each marker. This is valuable for studying admixture history, adaptive introgression, and mapping ancestry-associated traits. Input consists of a panel of ancestral allele frequencies and the admixed samples' read/genotype counts with genetic positions, along with a model of admixture pulses; output is per-site posterior ancestry probabilities. It is built from source (the russcd/Ancestry_HMM GitHub repository) against the Armadillo C++ linear-algebra library, provided in a CentOS 7 + Armadillo/OpenBLAS container. | Computational genomics, Population genetics | population-genetics | Documentation, Uses, and more | |
| angsd | mcc, lcc | 1.0 | ANGSD (Analysis of Next Generation Sequencing Data) is a C/C++ program specialized for population-genetic inference that works directly from genotype likelihoods rather than requiring hard genotype calls, making it particularly powerful for low- and medium-coverage data where genotype uncertainty is high. Reading BAM/CRAM alignments (or existing genotype-likelihood files), it estimates the site-frequency spectrum, per-site allele frequencies, nucleotide diversity and neutrality statistics (Watterson's theta, Tajima's D), Fst and population differentiation, admixture and PCA inputs (via genotype-likelihood covariance for tools like PCAngsd/NGSadmix), association tests, and inbreeding/kinship, and it propagates uncertainty through these estimates. Analyses are selected through a system of options controlling major/minor allele determination, allele-frequency estimation, site-allele-frequency likelihoods, the genotype-likelihood model, and filtering, and are often followed by realSFS to fold or optimize the SFS. This environment pins version 0.931 and is provided as a conda app in a CentOS 8 baseline bioinformatics/cheminformatics container. | Ngs Data Analysis, Genome-Wide Association Studies, Population Genomics, Genetic Variation | Biology, Biological Sciences | Application | Documentation, Uses, and more |
| annoy | lcc | 1.0 | Annoy (Approximate Nearest Neighbors Oh Yeah, conda-forge python-annoy 1.17.0) is a C++ library with Python bindings, originally developed at Spotify, for fast approximate nearest-neighbor (ANN) search over collections of high-dimensional vectors. It builds a forest of randomized projection (hyperplane-splitting) trees over the vectors and supports several distance metrics — Euclidean, angular (cosine), Manhattan, Hamming, and dot product — trading a small amount of accuracy for large gains in query speed and memory efficiency. A distinctive feature is that its indexes are memory-mapped from disk, so an index can be built once, saved, and then shared read-only across many processes with minimal RAM overhead, which is useful for serving similarity queries at scale. In Python one creates an index object for a given dimension and metric, adds vectors to it, builds a chosen number of trees, optionally saves and reloads the index, and queries for nearest neighbors by item or by vector. This environment also provides numpy, scipy, pandas, matplotlib, and h5py, so it fits recommendation, embedding-lookup, and clustering/dimensionality-reduction workflows (e.g., neighbor graphs for single-cell or NLP embeddings). | Machine Learning, Data Science, Approximate Nearest Neighbors, High-Dimensional Data | Machine Learning, Data Mining | Library | Documentation, Uses, and more |
| ansysedt | lcc | N/A | ANSYS Electronics Desktop (AEDT) 23R2, the integrated environment for electromagnetic, circuit, and electro-thermal simulation (including HFSS, Maxwell, Q3D, and Icepak workflows). Used by engineers for antenna, RF/microwave, signal/power integrity, and electric machine design and analysis. | Electrical Engineering, Computational Electromagnetics | Electromagnetic Simulation Software | Documentation, Uses, and more | |
| antismash | mcc | 1.0 | antiSMASH (antibiotics & Secondary Metabolite Analysis Shell), version 8.0.0, is the leading tool for genome-wide identification, annotation, and analysis of secondary-metabolite biosynthetic gene clusters (BGCs) in bacterial and fungal genomes. Using curated profile-HMM rule sets it detects a very broad range of cluster classes — polyketide synthases (PKS), non-ribosomal peptide synthetases (NRPS), terpenes, RiPPs, siderophores, and many more — and then enriches each cluster with detailed annotations: predicted substrate specificities, domain architectures, comparison to known clusters in the MIBiG database (ClusterBlast/KnownClusterBlast), and gene functional assignments. Input is an annotated genome (GenBank/EMBL) or a FASTA sequence; output is an interactive, self-contained HTML report plus GenBank and JSON result files. It can be run from the command line or via its public web server. antiSMASH is central to natural-product discovery, metabolic-potential surveys, and comparative genomics of biosynthetic pathways. | Bioinformatics, Genome Mining, Biosynthetic Gene Clusters, Secondary Metabolites | Bioinformatics, Biological Sciences | Annotation Tool | Documentation, Uses, and more |
| anvio | mcc, lcc | 1.0 | Anvi'o (installed via pip from the v7.1 release with a bioconda dependency stack) is an integrated, community-driven analysis-and-visualization platform for microbial 'omics, especially metagenomics, and is best known for its interactive, browser-based interface for exploring complex datasets. It organizes data around a 'contigs database' and 'profile databases' that store gene calls, functional and taxonomic annotations, coverage, and single-nucleotide variant information across samples, enabling human-guided and automatic genome binning (recovery of metagenome-assembled genomes), refinement of bins, and quality assessment. Beyond binning, it supports pangenomics (comparing gene content across genomes), phylogenomics, metabolism/pathway prediction, single-amino-acid and single-nucleotide variant profiling, and rich interactive visualizations of coverage and phylogenetic relationships. A typical workflow generates a contigs database from an assembly, runs HMM and annotation steps, profiles each sample against the mapped BAMs, merges the samples, and opens the interactive interface to explore and refine bins. This container bundles anvi'o alongside supporting tools such as BUSCO, Bowtie2, MetaBAT2, and fineSTRUCTURE. | Omics Data Analysis, Microbiome, Metagenomics, Bioinformatics | Biological Sciences | Bioinformatics Tool | Documentation, Uses, and more |
| aocc | mcc | N/A | AOCC a set of production compilers optimized for software performance when running on AMD host processors using the AMD “Zen” core architecture.The AOCC compiler environment simplifies and accelerates development and tuning for x86 applications built with C, C++, and Fortran languages. Description Source: https://www.amd.com/en/developer/aocc.html |
Compiler, C, C++, Fortran, Optimization | Software Engineering, Computer & Information Sciences | Service | Documentation, Uses, and more |
| aocl-sparse | mcc | N/A | AOCL-Sparse contains basic linear algebra subroutines for sparse matrices and vectors optimized for AMD processors. In addition, AOCL-Sparse includes iterative sparse solvers for solving linear system of equations. It is designed to be used with C and C++. | Sparse Matrix Operations, Fpga Optimization, Opencl Applications, High-Performance Computing | Mathematics, Engineering & Technology | Library | Documentation, Uses, and more |
| apptainer | ecc | N/A | Apptainer, formerly known as Singularity, is an open-source container platform designed to create portable and reproducible environments for scientific computing. It allows users to package applications, their dependencies, and data in a single container that can be run consistently across different computing environments. Description Source: https://apptainer.org/ |
Containerization, Application Management, Devops, Software Development | Computer Science, Engineering & Technology | Development | Documentation, Uses, and more |
| argweaver | mcc | 1.0 | ARGweaver infers the ancestral recombination graph (ARG) — the set of coalescent genealogies and recombination breakpoints along a genome — from population-scale DNA sequence data, using a Bayesian Markov chain Monte Carlo sampler under the sequentially Markov coalescent. The sampled ARGs encode the full local genealogical history of a sample of sequences and allow inference of quantities such as local time to most recent common ancestor, recombination rates, and demographic and selective history. Inputs are aligned sequences (e.g., in its .sites format) together with mutation- and recombination-rate parameters and a discretized set of coalescent time points; its main sampler program runs the MCMC and emits sampled ARGs (.smc files), with companion utilities to summarize and extract statistics from the samples. This is version 20200306 installed from the genomedk conda channel. | bioinformatics, genetics, population genetics, software tool | Computational Evolutionary Biology, Bioinformatics | Command-line tool | Documentation, Uses, and more |
| arrayfire | lcc | N/A | Cell ranger tools . | Software Library, Parallel Computing, Gpu Acceleration, Linear Algebra, Signal Processing, Image Processing, Statistics | Parallel Computing, Computer & Information Sciences | Library | Documentation, Uses, and more |
| ase | lcc | N/A | ASE is an Atomic Simulation Environment written in the Python programming language with the aim of setting up, steering, and analyzing atomistic simulations. Description Source: https://wiki.fysik.dtu.dk/ase/index.html |
Computational Software, Python Library, Molecular Simulations | Chemistry, Other Natural Sciences | Tool | Documentation, Uses, and more |
| assembly-stats | mcc | 1.0 | assembly-stats (bioconda 1.0.1) is a small, fast command-line utility that computes summary statistics for genome-assembly and sequence files in FASTA or FASTQ format. For a given set of sequences it reports total assembly length, the number of sequences (contigs/scaffolds), minimum/maximum/mean/median sequence length, N50 and L50 (and related Nx/Lx metrics), N-count, and GC content, giving a quick quantitative picture of assembly contiguity and composition. It handles gzip-compressed input and can process multiple files, producing either a human-readable summary or a tab-delimited table convenient for comparing many assemblies at once. It is commonly used to benchmark and compare the output of different assemblers or successive curation steps (e.g., before and after purge_dups or scaffolding). | Genome Assembly, Bioinformatics | Biology, Biological Sciences | Stand-Alone Tool | Documentation, Uses, and more |
| astra-toolbox | ecc | N/A | ASTRA Toolbox 2.3.0 installed via Conda environment. | tomography, image reconstruction, medical imaging, simulation | Image Reconstruction, Tomography | Library | Documentation, Uses, and more |
| astral | mcc | 1.0 | ASTRAL estimates a species tree from a collection of unrooted gene trees under the multi-species coalescent (MSC) model, which explicitly accounts for gene-tree/species-tree discordance caused by incomplete lineage sorting (ILS). Rather than concatenating alignments, it takes previously inferred individual gene trees as input and finds the species tree that maximizes the number of shared induced quartet topologies, a statistically consistent estimator under the coalescent. It is fast enough for hundreds to thousands of taxa and genes and reports branch support as local posterior probabilities along with coalescent (internal) branch lengths in coalescent units. Input is a file of Newick gene trees (optionally with a mapping of individuals to species); output is the estimated species tree with support annotations. This is version 5.7.3 installed from the phylofisher conda channel. | Bioinformatics, Phylogenetics | Phylogenomics inference tool | Documentation, Uses, and more | |
| at-spi2-atk | mcc | N/A | Applications that provide accessibility through the ATK interfaces need a way to translate those interfaces to AT-SPI2 DBus calls. This module, at-spi2-atk, provides that translation bridge. Description Source: https://gitlab.gnome.org/Archive/at-spi2-atk |
Assistive Technology, Accessibility, User Interface | General, Computer & Information Sciences | Interface | Documentation, Uses, and more |
| at-spi2-core | mcc | N/A | AT-SPI2-Core is the central component of the Assistive Technology Service Provider Interface, providing APIs for assistive technologies to interact with Linux desktop applications, thereby enhancing accessibility for users with disabilities. | Accessibility, Assistive Technology, Gui Applications | General, Computer Science | Library | Documentation, Uses, and more |
| atacseqqc | mcc | 1.0 | ATACseqQC is a Bioconductor R package dedicated to quality control of ATAC-seq (Assay for Transposase-Accessible Chromatin) experiments, helping assess whether a library is suitable for downstream chromatin-accessibility analysis. It computes diagnostic metrics and plots including the fragment-size distribution (revealing the characteristic nucleosome-free, mono-, and di-nucleosome periodicity), library-complexity/duplication estimates, transcription-start-site (TSS) enrichment, and it can shift reads by the standard Tn5 +4/-5 offset and split them into nucleosome-free and nucleosome-bound fractions for footprinting. It operates in R on aligned BAM files with genome/annotation packages, via functions such as fragSizeDist, readBamFile, splitGAlignmentsByCut, and factorFootprints. This is version 1.24.0. | Documentation, Uses, and more | |||
| atk | mcc | N/A | ATK provides the set of accessibility interfaces that are implemented by other toolkits and applications. Using the ATK interfaces, accessibility tools have full access to view and control running applications. Description Source: https://hprc.tamu.edu/software/aces/ |
Computational Chemistry, Materials Science, Quantum Mechanics, Molecular Dynamics | Physical Sciences | Simulation Software | Documentation, Uses, and more |
| atomsk | lcc | 1.0 | Atomsk is a command-line program for creating, converting, and manipulating atomic-scale data files used in materials science and molecular-dynamics/ab-initio simulations. It can build model systems from scratch — perfect crystals from lattice parameters, supercells, nanowires, nanotubes, and polycrystals — and apply a wide range of transformations such as introducing dislocations, point defects, grain boundaries, strains, and rotations. A major strength is format interconversion: it reads and writes many simulation formats (LAMMPS data/dump, VASP POSCAR, XSF, CFG, XYZ, Quantum ESPRESSO, and more), acting as a universal translator between codes. Operations are composed as a sequence of options, for example creating a crystal in a chosen lattice and then applying mode and option flags, and it supports batch and one-in/one-out modes. This is version 0.13.1 from conda-forge. | atomic structures, materials science, simulation, file conversion | Computational Materials Science, Atomic and Molecular Physics | Command-line tool | Documentation, Uses, and more |
| augustus | lcc | N/A | AUGUSTUS is a program that predicts genes in eukaryotic genomic sequences. It can be run on this web server, on a new web server for larger input files or be downloaded and run locally. Description Source: https://bioinf.uni-greifswald.de/augustus/ |
Gene Prediction, Eukaryotes, Bioinformatics | Bioinformatics, Biological Sciences | Bioinformatics Tool | Documentation, Uses, and more |
| augustus-braker | lcc | N/A | AUGUSTUS with BRAKER is a genome annotation pipeline for eukaryotic gene prediction. BRAKER combines AUGUSTUS and GeneMark with RNA-seq and/or protein evidence to train gene-prediction models and produce structural gene annotations. Used in genomics for annotating newly assembled genomes. Provided as a Conda environment. | genome annotation, gene prediction, bioinformatics | Bioinformatics | Genome annotation software | Documentation, Uses, and more |
| autoconf | mcc, lcc, ecc | N/A | Autoconf is an extensible package of M4 macros that produce shell scripts to automatically configure software source code packages. These scripts can adapt the packages to many kinds of UNIX-like systems without manual user intervention. Autoconf creates a configuration script for a package from a template file that lists the operating system features that the package can use, in the form of M4 macro calls. Description Source: https://www.gnu.org/software/autoconf/ |
Build Automation, Software Configuration, Makefiles | General, Engineering & Technology | Build Tools | Documentation, Uses, and more |
| autoconf-archive | mcc, lcc, ecc | N/A | The GNU Autoconf Archive is a collection of more than 500 macros for GNU Autoconf that have been contributed as free software by friendly supporters of the cause from all over the Internet. Every single one of those macros can be re-used without imposing any restrictions whatsoever on the licensing of the generated configure script. Description Source: https://github.com/autoconf-archive/autoconf-archive |
autoconf, macros, configuration, software development | General, Software Engineering | Service | Documentation, Uses, and more |
| autodock-vina | mcc | 1.0 | AutoDock Vina is a widely used open-source program for molecular docking and structure-based virtual screening, predicting the bound conformation and binding affinity of a small-molecule ligand within a protein's binding site. It uses an efficient gradient-based local-optimization search combined with an empirical scoring function, making it substantially faster and easier to use than the classic AutoDock4 while giving competitive accuracy. Inputs are the receptor and ligand in PDBQT format (prepared with AutoDockTools/MGLTools, which assign charges and rotatable bonds) and a defined search box specified by center and size coordinates; output is a set of ranked docked poses with predicted affinities in kcal/mol. This is version 1.1.2, a mainstay of computational drug discovery and virtual-screening campaigns. | Molecular Docking, Bioinformatics, Computational Chemistry, Drug Discovery | Biology, Molecular Docking | Application | Documentation, Uses, and more |
| automake | mcc, lcc, ecc | N/A | autoamke software. | Build Automation, Software Development, Compilers | General, Engineering & Technology | Compiler | Documentation, Uses, and more |
| autotools | mcc, lcc, ecc | N/A | Developer utilities | Programming, Software Development, Build Automation | Software Engineering, Systems & Development, Computer & Information Sciences | Build Automation | Documentation, Uses, and more |
| aws-cli | mcc | 1.0 | The AWS CLI is Amazon's unified command-line client for programmatically managing and interacting with the full range of Amazon Web Services from a terminal or script. It exposes services such as S3 object storage, EC2 compute, IAM, and hundreds of others through a consistent per-service, per-operation command grammar, with credentials and defaults supplied via a configuration step, environment variables, or IAM roles. In HPC contexts it is most often used for bulk data transfer to and from S3, such as copying or syncing objects to stage large datasets or archive results, and it supports output formatting (JSON/text/table), profiles, pagination, and presigned URLs. This is the version 2 release installed from the official AWS CLI v2 bundled installer, which ships its own embedded Python runtime. | Aws, Command Line Interface, Cloud Management | General, Computer & Information Sciences, Engineering & Technology | Command Line Tool | Documentation, Uses, and more |
| aws-snowball | mcc | N/A | snowball software. | Data Transfer, Cloud Computing, AWS | Data Management, Computer Science | Data Transfer Service | Documentation, Uses, and more |
| azcopy | mcc | 1.0 | AzCopy is Microsoft's high-performance command-line utility for copying data to and from Azure Storage, specifically Azure Blob and Azure File services, optimized for large, high-throughput transfers with concurrency, resumable jobs, and integrity checks. It authenticates via Azure AD login or shared-access-signature (SAS) tokens embedded in the URL, and supports uploading local files/directories, downloading, and server-to-server copies between storage accounts without staging data locally. Its core operations copy a source to a destination (optionally recursively) and sync a local directory to a container, with flags controlling concurrency, blob tier, and include/exclude patterns, while a jobs facility inspects and resumes transfers. In an HPC context it is the standard tool for moving large datasets between cluster storage and Azure cloud object storage. This is the v10 Linux binary. | Data Transfer, Cloud Storage, Azure, Command Line Tool | Computer Science, Data Engineering | Command Line Utility | Documentation, Uses, and more |
| bacmet | mcc | 1.0 | BacMet is a curated, experimentally confirmed database of genes conferring resistance to antibacterial biocides (antiseptics, disinfectants, preservatives) and to heavy metals, used to screen bacterial genomes and metagenomes for the genetic basis of biocide/metal resistance and potential co-selection of antibiotic resistance. This environment does not add a bespoke BacMet program; instead it provides NCBI BLAST+ and DIAMOND as the search engines to query nucleotide or protein sequences against the BacMet reference sequences, which are staged locally at /share/singularity/data/bacmet. A typical workflow builds or uses a prebuilt database from the BacMet FASTA (experimentally confirmed or the larger predicted set) and runs a fast DIAMOND protein search (for large protein sets) or a sensitive NCBI BLAST search (for nucleotide or sensitive queries) with a chosen identity/coverage cutoff, then filters hits to annotate resistance genes and their target compounds. It is commonly applied in environmental microbiology and AMR surveillance to link resistance gene content to biocide/metal exposure. | Documentation, Uses, and more | |||
| bader | mcc | 1.0 | Bader is the Henkelman group's implementation of Bader charge analysis (the 'atoms in molecules' / quantum theory of atoms in molecules approach), version 1.0.5, used in computational chemistry and materials science to partition electronic charge among atoms. It reads a charge-density grid — most commonly a VASP CHGCAR/AECCAR file, but also Gaussian cube format — and partitions space into atomic Bader volumes bounded by the zero-flux surfaces of the charge density, then integrates the density within each volume to assign a net charge to every atom. This yields quantitative oxidation-state and charge-transfer information without reference to arbitrary basis-set partitioning. A typical VASP workflow sums the core and valence densities and runs the analysis against that summed reference density, producing ACF.dat (per-atom charge and volume), BCF.dat, and AVF.dat output files. It is a small, fast, standalone Fortran executable widely cited in surface-science and catalysis studies. | Documentation, Uses, and more | |||
| bakta | mcc, lcc | 1.0 | Bakta is a tool for rapid, standardized structural and functional annotation of bacterial genomes, metagenome-assembled genomes, and plasmids, producing output suitable for direct database submission. Structurally it predicts protein-coding genes (Prodigal), tRNAs, tmRNA, rRNAs, ncRNAs and CRISPR arrays; functionally it assigns products by combining a compact database of pre-computed, unique protein sequences (UPS/IPS) with homology and profile searches, yielding consistent gene identifiers, dbxref cross-references, and gene symbols. It takes an assembly FASTA and writes richly annotated GFF3, GenBank, EMBL, FASTA (nucleotide and protein), and TSV feature tables. This is version 1.11.0 and requires a downloaded Bakta reference database. | bacterial genome annotation, prokaryotic genomes, bioinformatics, functional annotation, gene prediction, comparative genomics, sequence analysis | Microbial genomics, Bioinformatics | Genome annotation software | Documentation, Uses, and more |
| bam-readcount | lcc | 1.0 | bam-readcount (bioconda 0.8) is a utility that generates detailed per-base sequencing metrics at specified genomic positions from a BAM alignment file, commonly used in variant validation and refinement. For each requested position (or region) it reports the reference base, the total read depth, and then, broken down by observed allele (each base plus insertions and deletions), the count of supporting reads along with summary statistics such as average base quality, average mapping quality, average distance from read end, strand distribution, and mismatch quality. This fine-grained allele-level evidence is widely used to filter and refine candidate somatic or germline variant calls (for example within the VarScan/mutation-validation workflows). Inputs are a coordinate-sorted, indexed BAM, the reference FASTA, and a list of positions or a BED region file; output is a tab-delimited line per position with the per-allele metric fields. Options let the user set minimum base/mapping quality thresholds and restrict to specific regions. | Bioinformatics, Genomics, Variant-Calling, Coverage-Assessment | Genomics, Biological Sciences | Command Line Tool | Documentation, Uses, and more |
| bash | ecc | N/A | Bash is a Unix shell and command language that is widely used for scripting and automation in Unix-like operating systems. | scripting, automation, command-line, Unix, Linux | Computer Science, Software Engineering | Shell | Documentation, Uses, and more |
| bayescan | mcc, lcc | 1.0 | BayeScan (version 2.1) identifies candidate loci under natural selection from population-genetic data by detecting F_ST outliers. It uses a Bayesian framework with a reversible-jump MCMC that decomposes locus-population F_ST coefficients into a locus-specific component (shared across populations, indicating selection) and a population-specific component (shared across loci, reflecting demography), then estimates the posterior probability that selection is acting at each locus. Loci with elevated locus effects are flagged as candidates for diversifying (positive) or balancing/purifying selection. Input is allele-count/genotype data in BayeScan's own format (convertible from tools like PGDSpider), and outputs include per-locus F_ST, posterior probabilities, q-values (for FDR control), and diagnostic plots. It is a standard method in landscape and population genomics for scanning genome-wide markers (SNPs, AFLPs) for signatures of local adaptation. | Genetic Analysis, Natural Selection, Bayesian Model Comparison | Genetics, Biological Sciences | Analysis Tool | Documentation, Uses, and more |
| bbmap | mcc | 1.0 | BBMap is the flagship aligner within BBTools, a suite of fast, memory-efficient, Java-based tools (written by Brian Bushnell at JGI) for processing high-throughput sequencing data. BBMap itself is a splice-aware, global short-read aligner that maps DNA/RNA reads to a reference with high accuracy even at high error or divergence, but the suite's most widely used component is bbduk.sh, an extremely fast tool for adapter and contaminant removal, quality trimming, k-mer-based filtering, and entropy filtering. Other utilities in the package include reformat.sh (format conversion and read manipulation), repair.sh (re-pairing disordered paired reads), dedupe.sh (duplicate removal), bbmerge.sh (merging overlapping paired reads), bbnorm.sh (k-mer coverage normalization), and clumpify.sh (reordering reads for better compression and duplicate marking). Each is a self-contained shell wrapper. This deployment installs BBMap version 39.08 from the prebuilt BBTools tarball. | Bioinformatics, DNA Sequencing, Sequence Alignment | Genomics, Biological Sciences | Bioinformatics Tool | Documentation, Uses, and more |
| bc | ecc | N/A | bc is an arbitrary precision calculator language that supports complex mathematical operations. | Calculator, Mathematics, Command-line tool | Applied Mathematics, Other Mathematics | Command-line tool | Documentation, Uses, and more |
| bcftools | mcc, lcc | 1.0 | BCFtools is the core companion toolkit to samtools/HTSlib for working with variant data in VCF and its binary BCF form. It provides a broad set of subcommands for calling variants from mpileup output, and for filtering, normalizing, annotating, merging, concatenating, subsetting, and computing statistics on variant files. Common operations include view (region/sample/type filtering), norm (left-align and split multiallelic sites), merge and concat (combining files), annotate (add or transfer INFO/ID fields), query (extract fields into custom tab-delimited output), and call for the modern consensus/multiallelic variant caller. It operates efficiently on bgzip-compressed, tabix-indexed files and streams via pipes, and its powerful include/exclude expression language lets users filter on any INFO, FORMAT, or genotype field. Version 1.18 is a standard, near-universal dependency in variant-calling and population-genomics pipelines. | Variant Calling, Genomics, Bioinformatics | Genomics, Biological Sciences | Bioinformatics Tool | Documentation, Uses, and more |
| bcftools-1.11-samtools | lcc | 1.0 | This is a paired environment providing BCFtools 1.11 together with SAMtools 1.11 so that alignment handling and variant handling are available at matching versions within one workflow. SAMtools manipulates SAM/BAM/CRAM sequence alignments, offering sorting, indexing, merging, filtering by region/flag/quality, statistics (flagstat, stats, depth), and pileup generation. BCFtools operates on VCF/BCF variant files and performs the mpileup+call variant-calling workflow, plus filtering, normalization (left-alignment and multiallelic splitting via norm), annotation, merging, concatenation, consensus-sequence generation, and per-sample statistics. Pinning both to 1.11 avoids incompatibilities between the alignment-manipulation and variant-calling stages of a pipeline. The two are provided as conda environments within a multi-app Miniconda container that also bundles molecular-dynamics tools, each exposed as an %apprun. | bioinformatics, genomics, variant analysis | Bioinformatics, Genomics | Command-line tool | Documentation, Uses, and more |
| bcl2fastq2 | mcc, lcc | 1.0 | Cell ranger tools . | bioinformatics, genomics, sequencing | Bioinformatics, Genomics | Command-line tool | Documentation, Uses, and more |
| bdftopcf | mcc | N/A | bdftopcf is a command-line font compiler for the X Window System that turns bitmap fonts in the text-based BDF format into the compact binary PCF format. PCF fonts are architecture-independent and load faster, which is why BDF fonts shipped with X are typically compiled with bdftopcf at install time. | font conversion, bitmap fonts, BDF, PCF, X11, command line, font utilities | "", Computer Science | Utility | Documentation, Uses, and more |
| beagle | mcc | 1.0 | Beagle is a widely used Java program for genotype phasing, genotype imputation, and identity-by-descent (IBD) inference from SNP genotype data in population and medical genetics. It phases unphased genotypes into haplotypes, imputes missing or ungenotyped markers using a reference haplotype panel (allowing low-density arrays to be filled to reference density), and can detect IBD/HBD segments. It reads and writes VCF files, with the reference panel supplied as a bref3/VCF and a genetic map used for accurate results. This build is version 5.2 (dated 21Apr21.304). | Phylogenetic Inference, Maximum Likelihood, Molecular Sequence Data, Parallel Processing, Likelihood Calculations | Genetics, Biological Sciences | Library | Documentation, Uses, and more |
| beagle-lib | lcc | N/A | BEAGLE is a high-performance library that can perform the core calculations at the heart of most Bayesian and Maximum Likelihood phylogenetics packages. It can make use of highly-parallel processors such as those in graphics cards (GPUs) found in many PCs. Description Source: https://github.com/beagle-dev/beagle-lib |
Computational Biology, Phylogenetic Inference, Likelihood Evaluation, Evolutionary Analysis, Bioinformatics | Biology | Library | Documentation, Uses, and more |
| beast | mcc, lcc | 1.0 | BEAST (Bayesian Evolutionary Analysis Sampling Trees), version 1 (the BEAST1 lineage), is a software package for Bayesian inference of evolutionary and phylogenetic parameters from molecular sequence data using Markov chain Monte Carlo (MCMC). Its particular strength is time-scaled phylogenetics: given dated sequences and a chosen molecular-clock model (strict or relaxed) and tree prior (coalescent models such as constant size, exponential growth, Bayesian skyline, or birth-death), it co-estimates the phylogeny, divergence times, substitution-model parameters, and demographic or phylogeographic parameters, integrating over uncertainty. It is heavily used in molecular epidemiology (e.g. estimating viral evolutionary rates and time to most recent common ancestor) and in dated species phylogenies. Analyses are configured in an XML input file (typically generated with the companion BEAUti GUI), run with the beast program (optionally BEAGLE-accelerated), and the resulting parameter and tree logs are summarized with tools such as Tracer, TreeAnnotator, and FigTree. This build is version 1.10.4 from Bioconda. | Bioinformatics, Computational Biology, Phylogenetics, Molecular Evolution, Bayesian Analysis | Systems Biology, Biological Sciences | Tool | Documentation, Uses, and more |
| beast2 | mcc, lcc | 1.0 | BEAST 2, version 2.7.7, is a cross-platform software package for Bayesian evolutionary analysis of molecular sequences using Markov chain Monte Carlo (MCMC), the successor rewrite of BEAST with a modular plugin architecture. It co-estimates time-calibrated phylogenies, substitution and molecular-clock models, and population/demographic parameters, and supports advanced models including relaxed clocks, birth-death and coalescent tree priors, tip-dating for measurably evolving populations (e.g. viral phylodynamics), and species-tree inference (StarBEAST). Analyses are defined in an XML file, usually generated with the companion BEAUti GUI, and then run through the BEAST engine; MCMC output (log and .trees files) is then assessed for convergence in Tracer, summarized with TreeAnnotator, and combined across runs with LogCombiner. This build links beagle-lib 4.0.1 to accelerate the phylogenetic likelihood calculations on CPU/GPU. It is a mainstay of molecular dating, phylogeography, and Bayesian phylodynamics. | Phylogenetics, Evolutionary Analysis, Molecular Sequences, Bayesian Inference | Ecology, Biological Sciences | Tool | Documentation, Uses, and more |
| bedops | mcc, lcc | 1.0 | BEDOPS (version 2.4.39) is a high-performance, memory-efficient toolkit for set-theoretic and statistical operations on genomic interval data, engineered to scale to whole-genome, whole-cohort datasets through a sorted-input streaming model. Core utilities include bedops for set operations (union, intersection, difference, symmetric difference, complement, and element-of), bedmap for mapping and computing statistics (counts, sums, means, overlap fractions) of one interval set onto another, closest-features for nearest-neighbor queries, and sort-bed for the fast lexicographic sorting the other tools require. It also ships numerous format converters (vcf2bed, gff2bed, gtf2bed, bam2bed, wig2bed, etc.) that normalize inputs into sorted BED. Because operations run in low, near-constant memory regardless of input size and compose via Unix pipes, BEDOPS is favored for very large ChIP-seq, epigenomics, and population-scale interval analyses. | Genomic Data Analysis, Bioinformatics, Computational Biology, Genomics, Data Analysis | Bioinformatics, Biological Sciences | Command-Line Tool | Documentation, Uses, and more |
| bedtools | mcc, lcc | 1.0 | BEDTools (version 2.31.1) is a widely used, high-performance toolkit for genome arithmetic — the manipulation and comparison of genomic interval data. It reads and writes common formats including BED, GFF/GTF, VCF, and BAM, and offers dozens of subcommands for set-theoretic and analytical operations: intersect finds overlaps between feature sets, merge collapses adjacent/overlapping intervals, and subtract, closest, complement, window, coverage, genomecov, flank, and slop among many others. It excels at tasks such as annotating variants against gene features, computing read depth over regions, and building or filtering peak sets. BEDTools follows a Unix-pipe philosophy, so commands stream and chain naturally. A companion Python interface (pybedtools) exists. Because of its speed and flexibility, it is a foundational utility in nearly every NGS analysis pipeline for ChIP-seq, RNA-seq, variant, and comparative-genomics work. | Bioinformatics, Genomics, Data Analysis, Tool | Biology, Biological Sciences | Application | Documentation, Uses, and more |
| bedtools2 | mcc | N/A | Collectively, the bedtools utilities are a swiss-army knife of tools for a wide-range of genomics analysis tasks. The most widely-used tools enable genome arithmetic: that is, set theory on the genome. Description Source: https://bedtools.readthedocs.io/en/latest/ |
Genomic Data Analysis, Bed Files, Bioinformatics | Biology, Biological Sciences | Tool | Documentation, Uses, and more |
| berkeley-db | mcc, lcc, ecc | N/A | Berkeley DB is a family of embedded key-value database libraries providing scalable high-performance data management services to applications. Description Source: https://docs.oracle.com/database/bdb181/ |
Database, Embedded Systems, High Performance, Data Management | Storage, Hpc Applications, Computer Science | Library | Documentation, Uses, and more |
| bigsi | mcc | 1.0 | BIGSI (BItsliced Genomic Signature Index) is a data structure and tool for indexing very large collections of bacterial and viral sequence datasets so that arbitrary k-mer or short-sequence queries can be searched across all of them rapidly. It represents each dataset as a Bloom filter of its k-mers, then stores these filters bit-sliced in a key-value database so that a query is answered by looking up only the bits corresponding to the query's k-mers across all samples simultaneously, scaling to hundreds of thousands of genomes. The typical workflow is to build per-sample Bloom filters (often from McCortex k-mer graphs), build the combined merged index, and then run sequence queries that return, for each sample, the percentage of query k-mers present — enabling searches for genes, plasmids, or resistance markers across whole sequence archives. This container is built from the upstream BIGSI Docker image with its Berkeley-DB 4.8.30 backend. | Documentation, Uses, and more | |||
| bin | mcc | N/A | Container | Documentation, Uses, and more | |||
| binutils | mcc, lcc | N/A | The GNU Binutils are a collection of binary tools. The main ones are: ld (the GNU linker), as (the GNU assembler), and gold (a new, faster, ELF only linker). Description Source: https://www.gnu.org/software/binutils/ |
Binary Tools, Linker, Assembler, Object Files | General, Computer & Information Sciences | System Software | Documentation, Uses, and more |
| bioawk | mcc | 1.0 | bioawk is an extension of Brian Kernighan's awk (bwk awk) that adds native understanding of common bioinformatics file formats, so users can manipulate sequence and annotation data with concise awk one-liners instead of custom parsers. By selecting a format name (fastx, sam, vcf, bed, gff, or a custom header), it exposes named fields such as name, seq, qual, chrom, and start, and it adds handy built-in functions like revcomp(), gc(), and length() for sequences. This makes tasks like computing GC content, extracting sequence names, filtering reads by length, or reformatting a GFF trivial. Version 1.0 is a lightweight but indispensable utility for quick command-line wrangling of FASTA/FASTQ/SAM/VCF/BED/GFF data in sequencing pipelines. | Bioinformatics, Data Processing, File Formats, Data Analysis | Bioinformatics, Biological Sciences | Scripting Tool | Documentation, Uses, and more |
| biobakery_workflows | lcc | 1.0 | bioBakery Workflows (version 3.0.0a7) is a collection of end-to-end meta'omics analysis pipelines that string together the individual bioBakery tools into reproducible, automated workflows. It provides ready-made pipelines for whole-metagenome shotgun data (running quality control with KneadData, taxonomic profiling with MetaPhlAn, and functional profiling with HUMAnN, then producing merged tables and summary visualizations) and for 16S amplicon data, as well as isolate assembly and metatranscriptome workflows. Built on the AnADAMA2 workflow engine, it manages task dependencies, parallel execution, and provenance, and can generate publication-style HTML/PDF visualization reports summarizing community composition and function across samples. Inputs are raw sequencing reads plus reference databases; outputs are the merged taxonomic/functional tables and report documents. It lets users run a complete, standardized microbiome analysis with a single command rather than orchestrating MetaPhlAn, HUMAnN, and KneadData separately. | bioinformatics, microbiome, data analysis, workflow | Computational Biology, Bioinformatics | Workflow Management | Documentation, Uses, and more |
| bioconductor | mcc | 1.0 | R | bioinformatics, genomics, R, data analysis | Bioinformatics, Biological Sciences | Library | Documentation, Uses, and more |
| bioconductor-champ | lcc | N/A | ChAMP (Chip Analysis Methylation Pipeline), a Bioconductor R package for end-to-end analysis of Illumina Infinium methylation BeadChip data (450K and EPIC). It loads IDAT files, performs quality control and probe filtering, several normalization methods, detection of differentially methylated positions (DMP) and regions (DMR via Probe Lasso, Bumphunter, DMRcate), copy-number inference, SVD-based batch-effect assessment, cell-type deconvolution, and GSEA. Used in epigenomics and biomedical DNA-methylation research. | Bioinformatics, Genomics, R | Bioinformatics, Molecular Biology | R Package | Documentation, Uses, and more |
| bioconductor-deseq2 | lcc | N/A | DESeq2 is a widely used Bioconductor/R package for differential gene-expression analysis of RNA-seq count data. It models counts with negative-binomial GLMs, performs normalization, dispersion estimation, and hypothesis testing, and is a standard tool in transcriptomics. | RNA-Seq, Differential Expression, Bioconductor, Genomics, Statistics | Genomics, Molecular Biology | R Package | Documentation, Uses, and more |
| bioconductor-dss | lcc | 1.0 | DSS (Dispersion Shrinkage for Sequencing) is a Bioconductor/R package for differential analysis of high-throughput sequencing count data, with particular strength in differential methylation from whole-genome or reduced-representation bisulfite sequencing (BS-seq). It models methylation counts with a Bayesian hierarchical beta-binomial (or, for RNA-seq, a gamma-Poisson) framework and uses shrinkage estimation of dispersion to gain power in small-sample designs, then detects differentially methylated loci (DML) and differentially methylated regions (DMR), supporting both two-group comparisons and general experimental designs via a linear-model interface. Typical usage in R constructs a BSseq data object, runs the DML test (including multi-factor variants), and then calls DML and DMR from those test results. This is version 2.38.0. | bioinformatics ,methylation, genomics, Bioconductor | Bioinformatics, Molecular Biology | R package | Documentation, Uses, and more |
| bioconductor-edger | lcc | N/A | edgeR is a Bioconductor/R package for differential expression analysis of count-based sequencing data (RNA-seq, ChIP-seq, etc.). It uses negative-binomial models with empirical Bayes dispersion estimation and exact/GLM tests, and is one of the most widely used transcriptomics analysis tools. | RNA-Seq, Differential Expression, Bioconductor, Statistics | Computational Biology, Bioinformatics | Statistical Analysis | Documentation, Uses, and more |
| bioconductor-lea | mcc | 1.0 | LEA (Landscape and Ecological Association studies) is a Bioconductor R package, here at version 3.18.0, for population-genomics analyses that combine genotype data with geographic and environmental information. It provides efficient functions for estimating individual ancestry coefficients and inferring the number of ancestral populations (via the snmf function, a fast sparse non-negative matrix factorization alternative to STRUCTURE/ADMIXTURE), imputing missing genotypes, and performing genotype-environment association (GEA) tests to identify loci potentially involved in local adaptation (via the lfmm/lfmm2 latent-factor mixed models that control for population structure). It reads common population-genetics formats through built-in converters (e.g., .geno, .lfmm, .ped, VCF) and returns ancestry matrices, cross-entropy criteria for model selection, and association statistics with correction for confounding. As an R/Bioconductor package it is loaded and driven programmatically; typical usage runs the snmf ancestry estimation to characterize structure and then the lfmm2 latent-factor tests to detect environmentally associated SNPs. It is widely used in landscape-genomics and local-adaptation studies. | Documentation, Uses, and more | |||
| bioconductor-maftools | mcc | 1.0 | maftools is a Bioconductor R package for the summarization, analysis, and visualization of somatic-variant data stored in Mutation Annotation Format (MAF), the standard output of cancer-genome mutation-calling pipelines such as those from TCGA. After reading one or more MAF files into a MAF object, it produces cohort-level summaries and rich visualizations including oncoplots (mutation waterfall matrices), transition/transversion spectra, lollipop plots of amino-acid changes on protein domains, and mutational-load and VAF distributions. Beyond description, it performs analytical tasks central to cancer genomics: detecting significantly co-occurring or mutually exclusive gene pairs, identifying likely driver genes with its oncodrive routine, enriching for known oncogenic pathways, extracting mutational signatures via NMF, comparing two cohorts, and integrating clinical annotations and copy-number/segmentation data. Version 2.22.0 is provided as a Bioconductor conda package in a Rocky 9 genomics container. | bioinformatics, cancer genomics, R package, data visualization | Cancer Research, Bioinformatics | R package | Documentation, Uses, and more |
| bionano | mcc | 1.0 | runBNG is a wrapper pipeline (AppliedBioinformatics/runBNG) that automates the common analysis tasks for BioNano Genomics optical genome mapping (OGM) data by orchestrating the underlying BioNano tools with a simpler command-line interface. Optical maps record the positions of sequence-specific labels along very long DNA molecules, providing long-range scaffolding information; runBNG streamlines de novo assembly of these molecule maps, alignment of maps, and — importantly — hybrid scaffolding that combines a BioNano optical-map assembly with a next-generation sequence assembly to produce longer, more contiguous scaffolds and to detect misassemblies. It is driven through subcommands (e.g. bng_assemble for the assembly step and bng_hybrid for the hybrid-scaffolding step) that take BioNano BNX/CMAP inputs and a sequence assembly, and it depends on R packages (data.table, igraph, intervals, argparser) for parts of its processing. It is used in genome-finishing projects to bridge gaps and validate scaffold order/orientation. | structural variation, genome assembly, optical mapping | Molecular Biology, Life Sciences | Bioinformatics Software | Documentation, Uses, and more |
| biopython | mcc | 1.0 | Biopython is a comprehensive collection of Python modules for computational molecular biology and bioinformatics, providing reusable objects and parsers so researchers can script analyses rather than build tooling from scratch. Its SeqIO and AlignIO modules read and write dozens of sequence and alignment formats (FASTA, FASTQ, GenBank, EMBL, Clustal, PHYLIP, etc.); Bio.Seq and Bio.SeqRecord model biological sequences with operations like transcription, translation, and reverse-complement; and further modules cover pairwise/multiple alignment, BLAST parsing, phylogenetics (Bio.Phylo), population genetics, 3D macromolecular structures (Bio.PDB), and programmatic access to NCBI Entrez and other online databases. It is used both interactively and as a library underpinning larger pipelines. Version 1.79 is a mature, foundational dependency across the Python bioinformatics ecosystem. | Biological Computation, Bioinformatics, Computational Biology, Molecular Biology, Structural Bioinformatics | Bioinformatics, Biological Sciences | Python Library | Documentation, Uses, and more |
| bis-snp | mcc | 1.0 | Bis-SNP (version 1.0.1 from bioconda) is a genotyping and DNA-methylation calling package for bisulfite-sequencing data that simultaneously infers genotypes (SNPs) and cytosine methylation levels while correctly accounting for the C-to-T conversion that bisulfite treatment introduces. Built on a GATK-style Bayesian framework, it distinguishes true genetic variants from methylation-driven base changes, which is essential for accurate methylation calling near polymorphic sites and for allele-specific methylation analysis. It takes a bisulfite-aligned, sorted/indexed BAM and a reference genome and outputs methylation calls (in CpG/CpH contexts) and genotype calls, commonly emitting per-cytosine methylation in VCF/BED-like formats for downstream differential-methylation analysis. It is a Java tool that runs its bisulfite genotyper on the aligned reads against the reference. It is used in whole-genome bisulfite sequencing (WGBS) and epigenetics studies where joint genotype/methylation accuracy matters. | bisulfite sequencing, DNA methylation, SNP calling, epigenomics, bioinformatics, NGS analysis | Epigenomics, Bioinformatics | Variant and DNA methylation analysis software | Documentation, Uses, and more |
| bismark | mcc, lcc | 1.0 | Bismark is a bisulfite-read aligner and methylation caller for whole-genome and reduced-representation bisulfite sequencing (WGBS/RRBS). It aligns bisulfite-converted reads to a bisulfite-converted reference (using Bowtie2 or HISAT2 as the underlying aligner) in a strand-aware manner and then determines the methylation state of every cytosine, distinguishing CpG, CHG, and CHH contexts. The workflow uses a genome-preparation step to build the converted reference indexes, the bismark aligner to map FASTQ reads and produce methylation-tagged BAMs, a deduplication step to remove PCR duplicates, and a methylation-extractor step to generate per-cytosine methylation calls, cytosine reports, and bedGraph/coverage files, along with summary HTML reports. It is a standard tool for DNA methylation studies. This build is version 0.25.1 from bioconda, paired with samtools 1.18 or newer for BAM handling. | Epigenetics, DNA Methylation, Bisulfite Sequencing, Genomics, Bioinformatics | Genomics, Biological Sciences | Alignment & Methylation Analysis | Documentation, Uses, and more |
| bison | mcc, lcc | N/A | Bison is a general-purpose parser generator that converts an annotated context-free grammar into a deterministic LR or generalized LR (GLR) parser employing LALR(1) parser tables. Description Source: https://www.gnu.org/software/bison/ |
Parser Generator, Programming Language Tool | General, Computer & Information Sciences | Parser Generator | Documentation, Uses, and more |
| blasr | mcc, lcc | 1.0 | BLASR (Basic Local Alignment with Successive Refinement), version 5.3.5, is PacBio's long-read aligner designed for the high error rates and long read lengths of SMRT sequencing. It maps single-molecule reads to a reference genome using a seed-and-refine strategy that clusters short exact-match anchors and then performs banded dynamic-programming refinement, making it suited to noisy CLR-era PacBio data. It accepts reads as FASTA/FASTQ, PacBio BAM, or the legacy HDF5/bax.h5 formats and can emit SAM/BAM or PacBio's tabular alignment formats. It offers options controlling seeding, best-N hits, and minimum alignment length and identity. Historically a core component of PacBio's resequencing and consensus workflows, it has largely been superseded by minimap2/pbmm2 for modern HiFi data but remains useful for legacy datasets and reproducing older pipelines. | Alignment, Sequencing, Genomics | Biology, Biological Sciences | Mapping | Documentation, Uses, and more |
| blast | mcc, lcc | 1.0 | NCBI BLAST+ is the Basic Local Alignment Search Tool suite for finding regions of local similarity between nucleotide and protein sequences and for searching sequences against databases. It provides the classic programs — blastn (nucleotide vs nucleotide), blastp (protein vs protein), blastx (translated nucleotide query vs protein db), tblastn, and tblastx — plus makeblastdb for building searchable databases and utilities like blastdbcmd for extracting sequences. Each search reports high-scoring segment pairs with alignment statistics (bit score, E-value, percent identity) and supports numerous output formats, most usefully the customizable tabular formats for downstream parsing. It is the foundational tool for sequence homology searching, annotation transfer, and database curation, with a typical workflow of building a database with makeblastdb and then searching it with one of the blast programs. This is version 2.12.0 from bioconda. | Bioinformatics, Sequence Alignment, Homology Search | Biological Sciences | Sequence Analysis Tool | Documentation, Uses, and more |
| blat | mcc, lcc | 1.0 | BLAT on DNA is designed to quickly find sequences of 95% and greater similarity of length 25 bases or more. It may miss more divergent or shorter sequence alignments. It will find perfect sequence matches of 20 bases. Description Source: https://genome.ucsc.edu/cgi-bin/hgBlat |
Sequence Analysis, DNA Alignment, Genomic Analysis | Biology, Biological Sciences | Tool | Documentation, Uses, and more |
| blobtoolkit | mcc | 1.0 | BlobToolKit (version 4.1.4 from the hcc channel) is a suite for quality control and decontamination of genome assemblies, generating and interactively exploring 'blob plots' that array contigs by GC content, read coverage, and taxonomic assignment. It ingests an assembly FASTA together with coverage from mapped reads (BAM), sequence-similarity hits (e.g., BLAST/Diamond against nt/UniProt), and BUSCO completeness results to build a BlobDir dataset, which can be filtered to flag or remove contaminant/foreign contigs and visualized in the BlobToolKit Viewer. Core stages include creating a BlobDir, adding coverage, hits, and BUSCO layers, filtering, and generating views such as snail, blob, and cumulative plots. It is a central component of reference-genome QC pipelines such as those at the Sanger/Darwin Tree of Life. Outputs support both publication figures and cleaned assemblies ready for submission. | genomics, bioinformatics, data visualization, software tools | Computational Biology, Bioinformatics | Analysis Tool | Documentation, Uses, and more |
| blobtools | mcc | 1.0 | BlobTools is a modular command-line and visualization tool for quality assessment and contamination screening of genome assemblies by combining three independent signals: GC content, read coverage, and taxonomic annotation of contigs. It creates 'blobplots' — scatter plots where each contig is a bubble positioned by GC (x) and coverage (y), sized by length, and colored by taxonomic assignment — making it easy to spot contaminant contigs (e.g. bacterial or symbiont sequence in a eukaryotic assembly) that cluster separately from the target organism. The workflow builds a BlobDB from an assembly FASTA, one or more BAM coverage files, and hit files from similarity searches (e.g. BLAST or DIAMOND against a taxonomically annotated database), then generates plots and tables and can extract subsets of reads/contigs by taxon; conceptually this proceeds through a create step to build the database, followed by view and plot steps to produce tables and blobplots. This is the DRL/blobtools v1.1.1 release with matplotlib/docopt/pysam/samtools dependencies from conda. | Genome Assembly, Data Visualization, Metagenomics | Biology, Biological Sciences | Bioinformatics Tool | Documentation, Uses, and more |
| boltz | lcc | 1.0 | Boltz (installed here as boltz[cuda] from PyPI, providing Boltz 2.2.1 with GPU support) is an open-source, AlphaFold3-class deep-learning model for predicting the 3D structure of biomolecular complexes, including proteins, nucleic acids (DNA/RNA), ligands, and their interactions, and it can additionally predict protein-ligand binding affinity. It takes an input specification of the complex (sequences and ligand definitions via a YAML/FASTA-style input, optionally with multiple-sequence-alignment inputs), runs diffusion-based structure generation on the GPU, and outputs predicted structures (mmCIF/PDB) with per-residue and interface confidence scores (such as pLDDT/PAE-style metrics) plus, for Boltz-2, affinity estimates. The prediction step can use an MSA server or a precomputed MSA, running on CUDA hardware for tractable runtimes. It gives academic users a freely usable alternative to closed structure-prediction services for complexes and screening. This build runs in a Python 3.12 virtual environment inside a GPU-enabled Rocky 8 container. | Documentation, Uses, and more | |||
| bonito | lcc | N/A | Bonito is Oxford Nanopore Technologies' deep-learning basecaller, provided on LCC as conda-backed modules (ccs/bonito-0.0.8 and ccs/bonito-0.3.2). It converts raw nanopore electrical signal (squiggle) into DNA/RNA sequence using neural-network models, producing high-accuracy reads for genomics workflows, and also supports model training and research basecalling, typically on GPUs. It is a real user-facing bioinformatics application used at the front of nanopore sequencing analysis pipelines. | bioinformatics, nanopore sequencing, basecalling, deep learning | Molecular Biology, Bioinformatics | Basecalling Software | Documentation, Uses, and more |
| boost | mcc, lcc | N/A | Boost free peer-reviewed portable C++ source libraries | C++ Libraries, Programming Support, Efficiency Enhancement | Computer Science, Software Engineering, Systems & Development, Computer & Information Sciences | Development Tool | Documentation, Uses, and more |
| bowtie | mcc, lcc | 1.0 | This is Bowtie2 version 2.3.5, a fast and memory-efficient aligner for mapping short DNA sequencing reads to reference genomes, using an FM-index (Burrows-Wheeler transform) so that even large mammalian genome indexes fit in modest memory. It supports gapped, local, and paired-end alignment, offers preset sensitivity/speed trade-offs from very-fast through very-sensitive in either end-to-end or local modes, reports multiple or best-effort alignments, and outputs standard SAM. The workflow first builds an index from the reference FASTA and then aligns single- or paired-end reads against that index. This particular build is the SRA-enabled distribution (bowtie2-2.3.5-sra-linux-x86_64), meaning it links NCBI's SRA toolkit so reads can be aligned directly from SRA accessions without a separate download step. It is provided as a prebuilt binary within a CentOS 8 Miniconda container of transcriptome/genome assembly and variant-analysis tools. | Bioinformatics, Genomics, DNA Sequencing, Sequence Alignment | Genetics, Biological Sciences | Alignment Tool | Documentation, Uses, and more |
| bowtie2 | mcc, lcc | 1.0 | Bowtie2 (version 2.4.4 from bioconda) is a fast, memory-efficient aligner for mapping DNA sequencing reads to reference genomes, using an FM-index (Burrows-Wheeler) to keep its genome index compact. It is particularly strong for reads roughly 50 bp and longer, supports gapped, local, and paired-end alignment, and offers presets trading speed against sensitivity (ranging from very-fast to very-sensitive, in end-to-end or local modes). Users first build an index from a reference FASTA with the companion bowtie2-build tool, then align single- or paired-end reads against it, producing SAM output that is typically piped to samtools for sorting/indexing. It is a foundational read mapper used across ChIP-seq, ATAC-seq, RNA-seq (as the engine under TopHat/for quantifiers), metagenomics, and general resequencing, and its index format underlies tools like Bismark and Kraken-adjacent pipelines. | Alignment, Sequencing Reads, DNA Sequences, High-Throughput Sequencing, Bioinformatics, Genomics | Bioinformatics, Biological Sciences | Sequence Alignment | Documentation, Uses, and more |
| bpp | mcc, lcc | 1.0 | BPP (Bayesian Phylogenetics and Phylogeography) is a program for inferring evolutionary relationships and demographic parameters under the multispecies coalescent (MSC) model directly from multilocus sequence alignments. It supports four analysis modes: estimating parameters (population sizes theta and divergence times tau) on a fixed species tree (A00), species delimitation on a guide tree (A10), joint species-tree inference and delimitation (A11), and species-tree estimation with an assignment of individuals to species (A01). Version 4 adds the MSC-with-introgression (MSci) model to accommodate gene flow. Input is a control file specifying the sequence data (phylip-style multilocus alignment), an Imap file mapping sequences to populations, priors, and MCMC settings; output includes posterior distributions of the parameters and, for delimitation runs, posterior probabilities of alternative species models. It is a mainstay of coalescent-based species delimitation and divergence-time studies. | phylogenetics, molecular evolution, comparative genomics | Evolutionary Biology, Life Sciences | Statistical Software | Documentation, Uses, and more |
| braker | lcc | 1.0 | BRAKER is an automated, fully unsupervised pipeline for structural annotation of eukaryotic genomes that predicts protein-coding gene structures by combining the self-training gene finder GeneMark-ES/ET/EP with AUGUSTUS. It trains AUGUSTUS on the fly using extrinsic evidence, and supports several evidence modes: RNA-seq alignments (BRAKER1), a large protein database via ProtHint/GeneMark-EP (BRAKER2), or both RNA-seq and protein evidence together (BRAKER3) for the most accurate results. Inputs are a soft-masked genome assembly plus aligned RNA-seq BAMs and/or a protein FASTA; the output is a GTF/GFF3 of predicted genes with coding sequences and proteins. This container bundles the full dependency stack — AUGUSTUS, GeneMark-ES, ProtHint, GenomeThreader, DIAMOND, and samtools/bcftools/htslib — which is the main practical hurdle to running BRAKER, making it a convenient one-stop annotation environment. | gene prediction, RNA-Seq, genomics, bioinformatics | Bioinformatics, Computational Biology | Bioinformatics Tool | Documentation, Uses, and more |
| braker3 | mcc | 1.0 | BRAKER3 (version 3.0.7.6, wrapped from the upstream TeamBRAKER Docker image) is a fully automated pipeline for structural genome annotation — predicting protein-coding gene structures in eukaryotic genomes. BRAKER3 integrates GeneMark-ETP and AUGUSTUS and uses both RNA-seq alignments and a protein database (e.g. OrthoDB) as extrinsic evidence, training the gene finders in an unsupervised, iterative manner and combining their predictions (via TSEBRA) into a consensus gene set. This tri-input strategy (genome + RNA-seq + proteins) generally yields higher-accuracy annotations than earlier BRAKER1/BRAKER2 modes that used only one evidence type. Inputs are a soft-masked genome FASTA plus RNA-seq BAM files and/or a protein FASTA; outputs are gene structures in GTF/GFF3 with predicted proteins and transcripts. It is a leading choice for annotating newly assembled non-model eukaryotic genomes. | Gene Prediction, Genome Annotation, Bioinformatics, RNA-Seq, Eukaryotic Genomes | Bioinformatics, Computational Biology | Bioinformatics Tool | Documentation, Uses, and more |
| breakdancer | mcc, lcc | 1.0 | BreakDancer (packaged from the upstream molecular/breakdancer Docker image) is a structural-variant (SV) detection tool that predicts genomic rearrangements from paired-end sequencing data by analyzing anomalies in read-pair mapping. It examines the mapping orientation and the deviation of insert sizes from the expected distribution to infer deletions, insertions, inversions, and both intra- and inter-chromosomal translocations, computing a confidence score for each predicted breakpoint. The primary executable, breakdancer-max, consumes a configuration file listing one or more BAM files and their library insert-size statistics (generated by the bundled bam2cfg.pl script from the BAMs), and emits a tab-delimited call set with breakpoint coordinates, SV type, size estimate, supporting read counts, and confidence score. The typical workflow first builds the configuration file from the BAMs and then runs the caller to produce the call set. As a read-pair-based caller it is best suited to Illumina paired-end data and is often used alongside split-read and read-depth methods for comprehensive SV discovery. | genomics, bioinformatics, structural variants, next-generation sequencing | Genomic Analysis, Bioinformatics | Analysis Tool | Documentation, Uses, and more |
| bsbolt | lcc | 1.0 | BSBolt (Bisulfite Sequencing Processing Toolkit) is an integrated tool for the analysis of bisulfite-sequencing data, covering whole-genome bisulfite sequencing (WGBS), reduced-representation (RRBS), and targeted assays. It provides a unified command-line interface for the full workflow: building a bisulfite-aware reference index, aligning bisulfite-converted reads (using a BWA-based, wildcard/conversion-aware alignment engine) to produce methylation-ready BAMs, and calling per-cytosine methylation to generate methylation matrices and CGmap/coverage outputs. It also includes a read-simulation module for generating synthetic bisulfite reads to benchmark methods. Its subcommands cover Index, Align, CallMethylation, and Simulate stages. It is a fast, single-package alternative to stitched-together bisulfite pipelines. This is version 1.4.8 installed from the cpfarrell conda channel. | bisulfite sequencing, DNA methylation, read alignment, variant calling, epigenomics, bioinformatics | Epigenomics, Bioinformatics | Analysis software | Documentation, Uses, and more |
| burai | lcc | N/A | BURAI is a graphical user interface for the Quantum ESPRESSO first-principles (DFT) electronic-structure package. It lets computational materials scientists and chemists build atomic models, configure SCF/relaxation/band-structure/DOS calculations, launch Quantum ESPRESSO runs, and visualize results (structures, charge densities, spectra) without hand-editing input files. It is commonly used for teaching and rapid setup of plane-wave DFT simulations. The module self-identifies as Category: Application. | quantum espresso, density functional theory, materials simulation | Computational Physics, Physical Sciences | Scientific Application | Documentation, Uses, and more |
| busco | mcc, lcc | 1.0 | BUSCO (Benchmarking Universal Single-Copy Orthologs), version 5.7.0, assesses the completeness of genome assemblies, annotated gene sets, and transcriptomes by measuring how many of a lineage's expected near-universal single-copy orthologs are recovered. It compares the input against OrthoDB-derived lineage datasets (e.g. bacteria_odb10, eukaryota_odb10, vertebrata_odb10) and reports each marker as Complete (single-copy or duplicated), Fragmented, or Missing, summarized in the familiar C/S/D/F/M string that serves as a standard assembly-quality metric. It runs in genome, transcriptome, or protein mode; genome mode uses gene predictors (Metaeuk or AUGUSTUS/MiniProt depending on configuration) plus HMMER/prodigal, while protein mode simply HMMER-searches supplied sequences, and an auto-lineage option can detect the appropriate dataset. Outputs include the short summary, full per-marker tables, and missing-marker lists, and BUSCO ships a companion plotting script for comparing runs. | Genomics, Genome Assembly, Gene Set, Transcriptome, Orthologs | Bioinformatics, Biological Sciences | Tools | Documentation, Uses, and more |
| bvalcalc | mcc | 1.0 | Bvalcalc (version 1.4.3 from conda-forge) computes B-values across a genome, where B quantifies the expected local reduction in neutral genetic diversity caused by background selection (linked selection against deleterious mutations). It models how the density and recombination environment of functional (e.g., coding/conserved) elements depress diversity at linked neutral sites, producing per-position or windowed B estimates that can be used to correct demographic and selection inferences. Inputs typically include annotations of functional regions and a recombination map plus mutation/selection parameters, and outputs are genome-wide B tracks suitable for downstream population-genetic modeling. It provides a fast, scriptable way to generate the background-selection expectations needed to distinguish the signature of positive selection from that of linked purifying selection. It is aimed at population geneticists studying diversity landscapes and selection. | Documentation, Uses, and more | |||
| bwa | mcc, lcc | 1.0 | BWA (Burrows-Wheeler Aligner), version 0.7.17, is a foundational software package for mapping DNA sequencing reads against a reference genome using an FM-index derived from the Burrows-Wheeler transform. It comprises three algorithms: BWA-backtrack for short (<~70 bp) reads, BWA-SW, and — most widely used today — BWA-MEM, which performs fast, accurate local alignment of reads from roughly 70 bp to a few megabases, handling paired-end data, split/chimeric alignments, and gaps well. The reference must first be indexed, after which reads are aligned with the MEM algorithm and the output typically piped into samtools for sorting and BAM conversion. Output is standard SAM/BAM with mapping-quality and alignment tags suitable for downstream variant calling (GATK, bcftools) and other analyses. Despite newer aligners, BWA-MEM remains a default, well-validated choice in resequencing and variant-discovery pipelines. | Sequence Alignment, Bioinformatics, Genomics, Next-Generation Sequencing (Ngs) | Biology, Biological Sciences | Bioinformatics Tool | Documentation, Uses, and more |
| bwa-mem2 | mcc | 1.0 | BWA-MEM2 is the optimized successor to the ubiquitous BWA-MEM short-read aligner, producing essentially identical alignments while running roughly 1.5-3x faster through SIMD acceleration and an improved FM-index. It maps DNA sequencing reads (typically Illumina, both single-end and paired-end, from ~70bp up to a few hundred bp) against a reference genome and emits SAM output suitable for sorting and variant calling. Usage mirrors BWA: an index is built from the reference, then reads are aligned against it, commonly piped straight into a samtools sort step. The main tradeoff is that its index is larger and index construction needs more memory than classic BWA, but at run time it is a drop-in replacement. It is a standard first step in genome resequencing, variant-discovery, and Hi-C scaffolding workflows. | bioinformatics, genomics, alignment, DNA sequencing | Bioinformatics, Genomics | Command-line tool | Documentation, Uses, and more |
| bwakit | mcc | 1.0 | bwa-kit is a self-contained distribution built around Heng Li's BWA aligner that bundles the aligner together with the helper scripts and third-party binaries needed for a complete, best-practice read-mapping and post-processing workflow, particularly for human resequencing. Beyond BWA-MEM alignment it includes utilities for generating an ALT-aware reference index and for correctly handling GRCh38 ALT contigs (postalt processing), plus bundled tools such as samtools, seqtk, and trimadap to sort, mark duplicates, and produce analysis-ready BAM/CRAM in one pass. Its run-bwamem script ties these steps together, taking a reference index and FASTQ reads and emitting sorted, ALT-processed alignments. It is a convenient way to get reproducible, ALT-aware human alignment. This is bioconda bwakit version 0.7.17.dev1. | bioinformatics, genomics, sequencing, data processing | Bioinformatics, Genomics | Command-line tool | Documentation, Uses, and more |
| bwameth | mcc | 1.0 | bwa-meth (bioconda 0.2.4) is a tool for aligning bisulfite-treated whole-genome bisulfite sequencing (WGBS) reads to a reference genome, wrapping BWA-MEM to handle the C-to-T (and G-to-A) conversions that bisulfite treatment introduces. It works by in-silico converting both the reference and the reads and mapping in the converted space, which lets it use BWA-MEM's accurate and fast alignment while correctly restoring original read/reference bases and strand information in the output, producing standard BAM files with methylation-aware tags suitable for downstream methylation calling (e.g., with MethylDackel). It is valued for being simple, fast, and accurate for directional WGBS libraries. The workflow first indexes the reference and then aligns the FASTQ reads against it, sorting the result into BAM. Inputs are FASTQ reads and a reference FASTA; the output is a sorted BAM of bisulfite alignments that feeds per-cytosine methylation extraction and downstream differential-methylation analysis. | DNA Methylation, Bisulfite Sequencing, Alignment, Bioinformatics | Genetics, Biological Sciences | Bioinformatics Tool | Documentation, Uses, and more |
| bzip2 | mcc, lcc, ecc | N/A | bzip2 is a freely available, patent free, high-quality data compressor. It typically compresses files to within 10% to 15% of the best available techniques (the PPM family of statistical compressors), whilst being around twice as fast at compression and six times faster at decompression. Description Source: https://sourceware.org/bzip2/ |
File Compression, Data Compression, Utility Software | General, Computer & Information Sciences | Compression Tool | Documentation, Uses, and more |
| c-blosc | mcc, lcc | N/A | c-blosc is a high-performance compressor optimized for binary data, designed to be used in conjunction with scientific computing applications. | compression, performance, binary data, scientific computing | Data Management, Data Compression | Library | Documentation, Uses, and more |
| ca-certificates-mozilla | ecc | N/A | Documentation, Uses, and more | ||||
| cactus | mcc, lcc | 1.0 | Progressive Cactus, version 3.2.1, is a reference-free whole-genome multiple-alignment and pangenome pipeline used in comparative genomics to align many genomes without designating any one as the reference. Guided by a user-supplied phylogenetic tree, cactus progressively aligns genomes and reconstructs ancestral sequences, producing a HAL (Hierarchical Alignment) file that captures the alignment plus inferred ancestors and can be queried/exported (to MAF, assembly-to-assembly alignments, etc.) with the HAL toolkit. The bundled cactus-pangenome subcommand builds graph pangenomes (using minigraph and the vg toolkit) from collections of samples, emitting GFA/vg graph and VCF outputs for pangenomic analysis. Input is a seqFile listing genome names, paths, and the guide tree; jobs are orchestrated with Toil for HPC/cloud scaling. This particular image extends the official CPU release with the unreleased cactus-panpatch wrapper and its panpatch binary for pangenome assembly patching. | Computational Software, Hpc Tools | Gravitational Physics, Physical Sciences | Simulation Software | Documentation, Uses, and more |
| cairo | mcc | N/A | Cairo is a 2D graphics library with support for multiple output devices. Currently supported output targets include the X Window System (via both Xlib and XCB), Quartz, Win32, image buffers, PostScript, PDF, and SVG file output. Description Source: https://cairographics.org/ |
Graphics Library, 2D Graphics, Rendering Engine | Graphics, Computer & Information Sciences | Library | Documentation, Uses, and more |
| calder | mcc | 1.0 | CALDER (Cancer Analysis of Longitudinal Data through Evolutionary Reconstruction) is a Java tool from the Raphael group that reconstructs the evolutionary history of a tumor from longitudinal bulk DNA sequencing—samples taken from the same patient at multiple time points. From variant/mutation cluster cellular-prevalence data across those timepoints, it infers a phylogenetic tree of tumor clones and their changing frequencies, exploiting the temporal ordering (a mutation present earlier constrains later evolution) to reduce ambiguity that confounds single-timepoint methods. It formulates the reconstruction as a constrained optimization solved with the GLPK linear-programming library (via GLPK-for-Java), and is packaged here as a built Java jar alongside an Absence-Aware-Clustering component. It is used in cancer-evolution research to trace clonal dynamics, treatment response, and the emergence of resistant subclones over time. | Documentation, Uses, and more | |||
| canu | mcc, lcc | 1.0 | canu software. | Genome Assembly, Nanopore Sequencing, Long-Read Sequencing, Bioinformatics | Genomics, Bioinformatics, Biological Sciences | Bioinformatics Tool | Documentation, Uses, and more |
| canu-2.2 | mcc | 1.0 | Canu is a de novo long-read genome assembler for PacBio and Oxford Nanopore data, implementing an overlap-layout-consensus strategy tuned for high per-base error rates. It runs three stages — read correction (using overlaps to correct the most reliable portions of raw reads), trimming (removing residual errors and chimeric junctions), and assembly (building a best-overlap graph and generating consensus contigs) — and can automatically parallelize across an HPC grid. It adaptively estimates parameters and handles genomes from small microbes to large eukaryotes, producing contiguous contigs plus an assembly graph. It can run as a single command over the reads with a genome-size estimate for PacBio, Nanopore, or PacBio-HiFi data, and the correction, trimming, and assembly stages can also be invoked independently. This is version 2.2 from bioconda, a later release than the 2.1.1 binary build also present in the catalog. | genome assembly, bioinformatics, high-throughput sequencing, single-molecule sequencing | Computational Biology, Bioinformatics | Assembly Software | Documentation, Uses, and more |
| caper | mcc | N/A | Caper is a tool for managing and executing workflows on cloud platforms, particularly designed for bioinformatics applications. | bioinformatics, workflow management, cloud computing, WDL | Bioinformatics, Biochemistry and Molecular Biology | Workflow Management Tool | Documentation, Uses, and more |
| ccl | mcc, lcc, ecc | N/A | CCL (Cpp Command Line Library) is a C++ library targeting Unix platform that makes it easy to define and parse command line arguments. It provides a simple and intuitive way to define and parse command-line arguments for C++ programs. | Command Line Interface, C++ Library | Documentation, Uses, and more | ||
| cd-hit | mcc, lcc | 1.0 | CD-HIT is a classic, highly efficient program for clustering and comparing large collections of protein or nucleotide sequences to reduce redundancy. It uses a greedy incremental clustering algorithm combined with short-word (k-mer) filtering to avoid the cost of all-versus-all alignment, grouping sequences that exceed a user-specified identity threshold and reporting a representative (the longest) for each cluster. The suite includes cd-hit for proteins and cd-hit-est for nucleotides, plus variants like cd-hit-2d (comparing two datasets) and cd-hit-est-2d. Inputs are FASTA files and outputs are a non-redundant representative FASTA plus a .clstr file describing cluster membership, with parameters setting the identity cutoff, word length, thread count, and memory limit. It is a standard step for building non-redundant databases and collapsing near-duplicate sequences; this is bioconda cd-hit 4.8.1. | Sequence Analysis, Clustering, Bioinformatics | Bioinformatics, Biological Sciences | Bioinformatics Software | Documentation, Uses, and more |
| cd-hit-auxtools | lcc | 1.0 | cd-hit-auxtools is a small collection of auxiliary command-line utilities distributed alongside the CD-HIT sequence-clustering package. Whereas the core CD-HIT tools cluster protein or nucleotide sequences by similarity to remove redundancy, the auxtools address related preprocessing tasks for next-generation sequencing data. The main programs are cd-hit-dup (identify and remove exact and near-duplicate reads from single- or paired-end FASTA/FASTQ, useful for collapsing PCR duplicates), cd-hit-lap (detect overlapping reads), and read-linker (used with cd-hit-dup to pair up read information). These utilities are typically used as a lightweight cleanup step ahead of assembly or clustering workflows. This build is version 4.8.1 from Bioconda, matching the CD-HIT 4.8.1 release. | bioinformatics, sequence analysis, clustering | Bioinformatics, Biological Sciences | Command-line tool | Documentation, Uses, and more |
| cdhit | lcc | 1.0 | CD-HIT (agbiome build 4.6.1) is a very widely used program for clustering and comparing large sets of biological sequences to reduce redundancy, available as cd-hit for proteins and cd-hit-est for nucleotide sequences. It uses a fast greedy incremental algorithm with a short-word (k-mer) filter: sequences are processed from longest to shortest, and each is either assigned to an existing cluster whose representative it exceeds a user-set identity threshold with, or becomes the representative of a new cluster. This makes it efficient enough to dereplicate millions of sequences, and it is commonly applied to build non-redundant reference databases, collapse near-identical reads or genes, and prepare inputs for downstream analyses. Key parameters include a sequence-identity threshold, a word length chosen to match that threshold, alignment-coverage cutoffs, and memory and thread controls. Inputs are FASTA; outputs are a FASTA of cluster representatives plus a .clstr file listing all members of each cluster. The package also includes utilities like cd-hit-2d and cd-hit-est-2d for comparing two datasets. | bioinformatics, sequence clustering, data reduction | Computational Biology, Biological Sciences | Command-line tool | Documentation, Uses, and more |
| cdna_cupcake | mcc, lcc | 1.0 | cDNA_Cupcake (version 22.0.0 from bioconda) is a set of Python scripts for post-processing full-length transcript data from PacBio Iso-Seq, cleaning and consolidating isoform models after alignment. Its central utility collapses redundant, genome-aligned full-length reads into a non-redundant set of unique transcript isoforms, and companion scripts filter isoforms by count/support, remove likely degraded 5' products, obtain per-isoform abundance, and prepare inputs for downstream tools such as SQANTI. It typically operates on minimap2/GMAP alignments (SAM/BAM) of HiFi/CCS Iso-Seq reads against a reference genome, producing collapsed GFF isoform models plus abundance and read-group files. A common step is the collapse of aligned reads by SAM into unique isoform models. It is a standard component of long-read transcriptome/isoform characterization workflows. | Documentation, Uses, and more | |||
| cdr-finder | mcc | 1.0 | CDR-Finder identifies Centromere Dip Regions, the characteristic hypomethylated valleys found within centromeric alpha satellite arrays, from long read sequencing data carrying base modification calls. It is a Snakemake workflow from the Eichler and Logsdon labs that takes an assembly, a BED file of centromeric regions and an alignment BAM with methylation tags, extracts the target sequence, derives per base methylation with modkit, averages it into fixed size windows, annotates alpha satellite with RepeatMasker, and then locates methylation valleys using height and prominence thresholds relative to the local median. Calls can be restricted to alpha satellite annotated sequence to suppress false positives in poorly curated regions, adjacent dips can be merged, edges extended, and thresholds adapted to centromeres whose overall methylation coverage is low. Outputs are a BED file of CDR coordinates per sample together with the windowed methylation signal, the repeat annotation used, and per contig plots showing the dips against the satellite structure. CDRs mark the site of the active kinetochore, so the tool is used in centromere and kinetochore studies built on telomere to telomere assemblies. | Documentation, Uses, and more | |||
| cell ranger | mcc | 1.0 | Cell Ranger is 10x Genomics' analysis pipeline suite for Chromium single-cell data, converting raw sequencer output into analysis-ready expression matrices. It performs demultiplexing of BCLs (the mkfastq stage), read alignment to a reference transcriptome, cell-barcode calling, UMI counting, and generation of filtered/raw feature-barcode matrices for single-cell gene expression (the count stage), plus aggregation of multiple libraries (the aggr stage) and support for Feature Barcoding, multiplexing, and V(D)J immune profiling. Outputs include the feature-barcode matrix (HDF5 and MatrixMarket formats), a cloupe file for the Loupe Browser, and an interactive HTML summary with QC metrics. It requires 10x-formatted FASTQs and a Cell Ranger-compatible reference package, and downstream matrices load directly into Seurat or Scanpy/Squidpy. This build is version 10.1.0 supplied as the official 10x tarball. | Documentation, Uses, and more | |||
| cellbender | mcc | 1.0 | CellBender is a single-cell RNA-seq preprocessing tool from the Broad Institute that uses a deep generative (variational autoencoder) model to remove technical artifacts from droplet-based (e.g. 10x) count matrices. Its principal module, remove-background, learns to distinguish genuine cell-associated transcripts from ambient RNA contamination—the soluble mRNA present in the droplet suspension that contaminates every barcode—and also flags empty droplets and corrects for barcode swapping, outputting a denoised count matrix and refined cell calls. It runs best on GPU and takes the raw (unfiltered) feature-barcode matrix from Cell Ranger as input. The cleaned matrix can then be loaded into Scanpy or Seurat, and removing ambient contamination typically sharpens clustering and marker-gene detection. This is version 0.3.0. | Single-Cell RNA-Seq, Data Processing, Denoising, Batch Effects, Bioinformatics | Genetics, Biological Sciences | Python Library | Documentation, Uses, and more |
| cellranger | mcc, lcc | 1.0 | Cell ranger tools . | Bioinformatics, Single-Cell RNA-Seq, Data Analysis | Biology, Biological Sciences | Documentation, Uses, and more | |
| cellranger-arc | lcc | N/A | Cell ranger arc tools . | Single-Cell Analysis, Chromatin Accessibility, Bioinformatics | Bioinformatics, Biological Sciences | Bioinformatics Tool | Documentation, Uses, and more |
| cellranger-atac | lcc | N/A | Cell ranger arc tools . | bioinformatics, single-cell genomics, ATAC-seq, chromatin accessibility | Biology, Genomics | Application | Documentation, Uses, and more |
| centrifuger | mcc | 1.0 | Centrifuger (bioconda 1.1.0) is a rapid, memory-efficient taxonomic classifier for metagenomic sequencing reads. It builds on the ideas of Centrifuge but uses a run-block-compressed Burrows-Wheeler transform (an FM-index variant) of the reference database, which greatly reduces the memory footprint required to index large, highly redundant collections of microbial genomes while retaining fast exact-match classification. Given reads, it assigns each to the most specific taxonomic node supported by its k-mer/substring matches and produces per-read classifications plus a Kraken-style abundance/summary report suitable for downstream profiling and visualization. It supports both short Illumina reads and long Nanopore/PacBio reads. The workflow is to first build an index from reference sequences and a taxonomy with its build step, then classify reads (single-end, paired-end, or long) against that index; an accompanying kreport tool converts output to a Kraken-format report. It is well suited to species/strain-level classification of large metagenomic datasets on memory-constrained systems. | Documentation, Uses, and more | |||
| cfour | mcc | N/A | cfour software. | Quantum Chemistry, Ab Initio Methods, High-Performance Computing | Chemical Sciences, Physical Chemistry | Scientific Software | Documentation, Uses, and more |
| cgal | mcc, lcc | N/A | CGAL is an open source software project that provides easy access to efficient and reliable geometric algorithms in the form of a C++ library. CGAL is used in various areas needing geometric computation, such as geographic information systems, computer aided design, molecular biology, medical imaging, computer graphics, and robotics. Description Source: https://www.cgal.org/ |
Computational Geometry, Geometric Algorithms, Software Library | Computer Science, Computer & Information Sciences | Algorithm Library | Documentation, Uses, and more |
| cgmaptools | lcc | 1.0 | cgmaptools is a command-line toolkit for the downstream analysis and visualization of DNA methylation data produced by whole-genome or reduced-representation bisulfite-sequencing pipelines, operating on the CGmap and ATCGmap file formats output by aligners such as BS-Seeker2 and BSMAP. It provides a broad set of subcommands for tasks such as extracting and converting methylation calls, computing per-context (CpG/CHG/CHH) methylation levels, binning methylation over genomic windows or genes, identifying differentially methylated regions between samples, computing methylation-level statistics, performing SNP calling from bisulfite data, and generating a variety of plots (fragment methylation distributions, metagene profiles, heatmaps, and dot/violin plots). It is designed to fit downstream of the read-mapping/methylation-calling step and to make methylome comparison and figure generation straightforward. This deployment is built from the guoweilong/cgmaptools GitHub repository within a conda environment that also provides R's optparse for its plotting scripts. | bioinformatics, genomics, methylation analysis | Bioinformatics, Biochemistry and Molecular Biology | Analysis Tool | Documentation, Uses, and more |
| cgns | mcc, lcc | N/A | The CFD General Notation System (CGNS) provides a standard for recording and recovering computer data associated with the numerical solution of fluid dynamics equations. Description Source: https://github.com/conda-forge/cgns-feedstock/blob/main/recipe/meta.yaml | Computational Software | Fluid & Plasma Physics, Physical Sciences | File Format | Documentation, Uses, and more |
| charliecloud | lcc | N/A | Lightweight user-defined software stacks for high-performance computing | Containerization, Software Development, Research Tools | High-Performance Computing, Computer & Information Sciences | Tool | Documentation, Uses, and more |
| charmm | lcc | N/A | CHARMM (Chemistry at HARvard Macromolecular Mechanics) version 44b1, a widely used molecular dynamics and molecular mechanics package for simulating biomolecular systems such as proteins, nucleic acids, lipids, and small molecules. Provides force fields, energy minimization, free-energy methods, and trajectory analysis for computational chemistry and structural biology research. | molecular dynamics, biomolecular simulations, computational chemistry, bioinformatics | Chemical Sciences, Physical Chemistry | Simulation Software | Documentation, Uses, and more |
| checkm-genome | mcc, lcc | 1.0 | CheckM is a widely used bioinformatics tool for assessing the quality of microbial genomes recovered from isolates, single cells, or metagenomes (metagenome-assembled genomes, MAGs). It estimates genome completeness and contamination by identifying sets of lineage-specific, collocated single-copy marker genes placed within a reference genome tree, then counting how many are present, missing, or duplicated. The primary workflow is the lineage_wf command, which runs tree placement, marker-set selection, HMMER-based marker identification, and quality-metric calculation in sequence, taking a directory of genome FASTA bins and producing a table of completeness, contamination, and strain-heterogeneity percentages. It also offers taxonomic-rank-specific marker sets (taxonomy_wf), plotting utilities (GC, coding density, nucleotide distribution), and coverage/bin-comparison tools. CheckM requires a downloaded reference data package and depends on HMMER, prodigal, and pplacer; this is bioconda checkm-genome 1.2.2. | Metagenomics, Microbial Ecology, Genome Quality Assessment | Environmental Biology, Biological Sciences | Genome Quality Assessment Tool | Documentation, Uses, and more |
| checkm2 | mcc | 1.0 | CheckM2 (version 1.1.0) rapidly estimates the completeness and contamination of microbial genomes, including isolate genomes and especially metagenome-assembled genomes (MAGs). Unlike its predecessor CheckM, which relied on lineage-specific marker sets, CheckM2 uses machine-learning models (gradient-boosted and neural-network predictors trained on genome features and KEGG/Pfam annotations via DIAMOND and Prodigal) to predict quality directly, giving accurate assessments even for organisms from novel or under-represented lineages and for reduced-genome symbionts. It reports percent completeness and percent contamination per genome in a summary table, which downstream tools use to filter MAGs to quality tiers (e.g. high-quality: >90% complete, <5% contaminated). Input is a directory of genome FASTA files and output is a quality-report TSV. It requires a downloaded reference database and is a standard QC step in genome-resolved metagenomics. | bioinformatics, genomics, metagenomics, quality control | Genomics, Microbiology | Command-line tool | Documentation, Uses, and more |
| checkv | mcc | 1.0 | CheckV is a tool for assessing the quality and completeness of viral genome sequences assembled from metagenomes or single-virus data, filling the same quality-control role for viruses that CheckM does for bacterial genomes. It compares input contigs against a reference database of complete viral genomes and a set of viral HMMs to estimate genome completeness, detect and trim host (proviral) flanking regions on integrated prophages, identify contamination, flag closed/complete genomes (e.g. by detecting direct terminal repeats), and assign each sequence to a quality tier (complete, high-, medium-, or low-quality, or not-determined) following MIUViG standards. Input is a FASTA of putative viral contigs (commonly from VirSorter2, geNomad, or similar); output is a set of tables reporting completeness, contamination, and provirus boundaries plus cleaned virus/proviral FASTA files. It offers an end-to-end mode that runs the full quality-assessment pipeline in one step. Version 1.0.3 is provided as a conda tool in a Rocky 9 + Miniconda genomics/metagenomics container. | bioinformatics, metagenomics, viral genomes, genome assembly | Bioinformatics, Biological Sciences | Command-line tool | Documentation, Uses, and more |
| chemshell | lcc | N/A | canu software. | Computational Chemistry, Quantum Mechanics, Molecular Dynamics | Computational Chemistry, Quantum Chemistry | Scientific Software | Documentation, Uses, and more |
| chewbbaca | lcc | 1.0 | chewBBACA (version 3.3.10) is a comprehensive suite for gene-by-gene (allele-based) bacterial typing, used to build and apply core-genome and whole-genome MLST (cgMLST/wgMLST) schemas from genome assemblies. Its modules cover the full workflow: CreateSchema defines a schema from a set of assemblies, AlleleCall assigns allele identifiers to each locus across genomes, and additional commands evaluate schema quality, detect paralogs, extract cgMLST loci at chosen presence thresholds, and annotate the schema. It uses Prodigal for gene prediction and BLAST-based comparisons, and produces allelic profile matrices that support high-resolution strain comparison, outbreak investigation, and phylogenetic clustering. Inputs are assembled genome FASTA files (and optionally an existing schema); outputs are per-locus allele FASTAs and a results allele-call matrix. It is a core tool in bacterial genomic epidemiology and public-health surveillance. | Whole-Genome Analysis, Microbial Genotyping, Bacterial Typing, Phylogenetic Inference, Bioinformatics | Genomics, Biological Sciences | Tool | Documentation, Uses, and more |
| chopper | mcc | 1.0 | Chopper is a fast command-line tool, written in Rust, for quality and length filtering (and optional trimming) of long-read sequencing data from Oxford Nanopore, effectively a high-performance reimplementation combining the functionality of the older Python tools NanoFilt and NanoLyse. It reads FASTQ from standard input and writes filtered FASTQ to standard output, so it slots naturally into a streaming pipeline. Options let the user set a minimum average read quality, minimum and maximum read length, trim a fixed number of bases from the read head and tail, and remove contaminant reads (such as lambda-phage control DNA) by mapping against a supplied reference. It is commonly used as the read-cleanup step at the front of Nanopore assembly and analysis workflows. This build is version 0.9.0 from Bioconda. | Bioinformatics, Ngs Data Analysis, Sequence Trimming, High-Throughput Sequencing | Bioinformatics, Biological Sciences | Bioinformatics Tool | Documentation, Uses, and more |
| ciao | lcc | N/A | Documentation, Uses, and more | ||||
| circos | lcc | N/A | Circos is a software package for visualizing data and information. It visualizes data in a circular layout — this makes Circos ideal for exploring relationships between objects or positions. Description Source: https://circos.ca/ |
Data Visualization, Genomics, Biological Sciences, Circular Layout | Bioinformatics, Biological Sciences | Visualization Tool | Documentation, Uses, and more |
| citcoms | lcc | N/A | CitcomS is a parallel finite-element simulation software used to model mantle convection and other large-scale geodynamic processes within the Earth. It is designed for high-performance computing environments and allows researchers to study thermal and compositional evolution of the mantle in spherical or regional geometries. | mantle convection, finite-element modeling, high-performance computing | Geodynamics, Earth Sciences | Simulation Software | Documentation, Uses, and more |
| clair3 | mcc | 1.0 | Clair3 is a deep-learning-based germline small-variant caller (SNPs and indels) developed for long-read sequencing, with support for short reads as well. It combines a fast pileup-based neural network for the majority of straightforward sites with a more computationally intensive full-alignment network for candidate difficult sites, achieving high accuracy on Oxford Nanopore and PacBio HiFi data where traditional callers struggle with platform-specific error profiles. It takes a coordinate-sorted, indexed BAM and a reference FASTA plus a platform-appropriate pretrained model, and outputs a VCF/gVCF of called variants; the model choice (ONT, HiFi, or Illumina) is a key parameter and this image bundles pretrained models. It is typically run via the run_clair3.sh wrapper, which is given the BAM, reference, output directory, sequencing platform, and model path. This container packages the upstream hkubal/clair3 v2.0.0 Docker image. | Variant Calling, Illumina, Sequencing Data | Bioinformatics, Biological Sciences | Variant Caller | Documentation, Uses, and more |
| clair3-rna | mcc | 1.0 | Documentation, Uses, and more | ||||
| clck | mcc, lcc | N/A | Intel Cluster Checker (clck) is a command-line diagnostic tool for assessing the health and readiness of an HPC cluster. It collects data from across the cluster's nodes and analyzes it to verify the system is correctly configured to run parallel applications, flagging problems such as hardware or configuration non-uniformity between nodes with graded severity levels, so administrators and users can pinpoint and resolve issues quickly. It ships as part of the Intel oneAPI HPC Toolkit. | cluster diagnostics, cluster health, system validation, Intel oneAPI, command-line | "", Informatics, Analytics and Information Science | Application | Documentation, Uses, and more |
| cli11 | lcc | N/A | CLI11 is a command line parser for C++11 and beyond, designed to be easy to use and flexible. | C++, Command Line Interface, Parsing, Library | Software Engineering, Other Computer and Information Sciences | Library | Documentation, Uses, and more |
| clumpp | mcc | 1.0 | CLUMPP (CLUster Matching and Permutation Program, prebuilt Linux 64-bit binary v1.1.2) addresses the label-switching and multimodality problems that arise when running population-structure programs such as STRUCTURE, ADMIXTURE, or TESS multiple times for a given number of clusters K. Because the cluster labels are arbitrary and can permute between replicate runs, CLUMPP finds the optimal alignment (permutation of cluster labels) across replicate membership-coefficient matrices so they can be sensibly averaged and compared. It offers three algorithms — FullSearch (exhaustive), Greedy, and LargeKGreedy — to search for the permutations that maximize similarity across runs, and it computes similarity statistics (H and H') quantifying agreement. Inputs are the Q-matrices from replicate runs (in its expected indfile/popfile format, often prepared by CLUMPAK or by hand); outputs are the permuted and averaged membership matrices plus a misc file with alignment statistics. Its averaged output is the standard input to visualization tools like distruct/pong for producing clean ancestry barplots. Here it is bundled alongside comparative-genomics and population-structure tools. | population structure, clustering alignment, genetic analysis | Life Sciences, Bioinformatics | Analysis Tool | Documentation, Uses, and more |
| clust | mcc | 1.0 | Clust is a tool for automatically extracting clusters of co-expressed genes from one or more gene-expression datasets, designed to find tight, biologically coherent clusters rather than force every gene into a group. Distinctively, it can integrate multiple datasets simultaneously—even across conditions, platforms, or species with a gene-mapping file—and returns only the genes that show consistent co-expression patterns, leaving out noisy or non-co-expressed genes. It automatically preprocesses the data (normalization, filtering) and optimizes cluster parameters, producing output clusters, their expression profiles as plots, and tables ready for downstream enrichment analysis. It takes a data directory and writes to a results directory, optionally with a normalization code and dataset-mapping file. Version 1.17.0 is used in transcriptomics to discover gene modules and candidate regulatory relationships across experimental conditions. | Clustering, Time Series Data, Pattern Identification | Other Computer & Information Sciences | Data Analysis Tool | Documentation, Uses, and more |
| cmake | mcc, lcc, ecc | N/A | CMake is an open-source, cross-platform family of tools designed to build, test and package software. | Build System, Cross-Platform, Software Development | Software Engineering, Systems & Development, Engineering & Technology | Developer Tools | Documentation, Uses, and more |
| cmdstanr 0.5.2 and opencl | lcc | 1.0 | CmdStanR is a lightweight, efficient R interface to CmdStan, the command-line interface to the Stan probabilistic programming language, used for full Bayesian statistical inference. It compiles Stan model files (.stan) into standalone executables and runs them from R, giving access to Stan's Hamiltonian Monte Carlo (NUTS) sampler, variational inference, and optimization, and returns draws that integrate with the posterior/bayesplot analysis ecosystem. This build (CmdStanR 0.5.2) additionally installs NVIDIA CUDA and OpenCL (ocl-icd-opencl-dev), enabling Stan's OpenCL backend to offload computationally heavy operations (such as large GLM likelihoods) to the GPU for faster sampling. A typical workflow compiles a Stan model into an executable and then samples from it across multiple chains, optionally enabling OpenCL via the appropriate compile/run flags. | Documentation, Uses, and more | |||
| code_aster | mcc | N/A | Code_Aster is an open-source software package for finite element analysis in structural mechanics, developed and maintained by EDF (Ãlectricité de France). It's widely used for simulating various engineering problems, including static and dynamic analyses, thermal behavior, and fluid-structure interactions. | Finite Element Analysis, Numerical Simulation, Open Source, Engineering | Engineering, Mechanical Engineering | Simulation Software | Documentation, Uses, and more |
| cogent | mcc | 1.0 | Cogent (from the Magdoll/Cogent repository) is a tool for reconstructing coding genomes and gene families from high-quality full-length transcript sequences — typically PacBio Iso-Seq HQ isoforms — when no reference genome is available. It partitions the input transcripts into gene families based on k-mer similarity, then reconstructs the minimal set of unique gene structures ("cogent contigs") that explain each family, effectively recovering the underlying coding sequence structure and collapsing redundant isoforms without a reference. This is valuable for non-model organisms where researchers have long-read transcriptomes but no assembled genome. Its environment bundles the supporting tools it uses: biopython, bx-python, the LP solver PuLP, the alignment library parasail-python, the long-read aligner minimap2, and Mash for sketch-based similarity. Inputs are full-length transcript FASTA/FASTQ files; outputs are the partitioned families and reconstructed gene models. It is used in long-read transcriptomics of reference-free species to build gene catalogs. | long-read sequencing, transcript reconstruction, isoform analysis | Life Sciences, Bioinformatics | Analysis Tool | Documentation, Uses, and more |
| colabfold | lcc | N/A | ColabFold is a protein structure prediction application that accelerates AlphaFold2 (and related models) by pairing it with fast MMseqs2-based multiple sequence alignment generation. It predicts 3D protein structures and complexes from amino acid sequences and is widely used in structural biology and computational biochemistry for rapid, high-accuracy modeling. | Protein Structure Prediction, Deep Learning, Biophysical Modeling | Bioinformatics, Biological Sciences | Protein Structure Prediction Tool | Documentation, Uses, and more |
| comfyui | ecc | 1.0 | ComfyUI is a node-based (graph) interface and execution engine for Stable Diffusion and other diffusion generative models, spanning text-to-image, image-to-image, inpainting, upscaling, and increasingly video generation. Instead of a fixed form, users wire together nodes representing checkpoints/model loaders, CLIP text encoders, samplers (a KSampler node with configurable steps, CFG scale, scheduler, and sampler type such as Euler or DPM++), VAE decoders, and post-processing, giving fine-grained control and reproducibility over each stage of the pipeline. It supports SD1.x/SD2.x/SDXL and many community checkpoints, LoRAs, ControlNet, and custom nodes, and caches unchanged parts of the graph so only downstream nodes re-execute on edits. In this deployment it runs headless as a CUDA 12.1 backend server that exposes an HTTP/WebSocket API, and it backs the Open OnDemand AI Chat / Image Studio apps rather than being used through its own browser GUI. Workflows are JSON documents that can be saved, shared, and submitted programmatically to the server for batch or app-driven generation. | Documentation, Uses, and more | |||
| compiler | mcc, lcc, ecc | N/A | A compiler is a special program that processes statements written in a particular programming language and turns them into machine language or "code" that a computer's processor uses. It typically acts as a translator that converts high-level programming languages into machine language. | Compiler, Software Development, Programming | Software Engineering, Computer & Information Sciences | Compiler | Documentation, Uses, and more |
| compiler-intel-llvm | ecc | N/A | LLVM is a free and open source compiler infrastructure originally built for C and C++. Although its name initially stood for low-level virtual machine, LLVM is now a technology that deals with much more than just virtual machines. The name is no longer officially an initialism. | llvm, high-performance computing, intel optimization | High Performance Computing, Software Engineering | Compiler | Documentation, Uses, and more |
| compiler-rt | mcc, lcc, ecc | N/A | builtins - low-level target-specific hooks required by code generation and other runtime components sanitizer runtimes - AddressSanitizer, ThreadSanitizer, UndefinedBehaviorSanitizer, MemorySanitizer, LeakSanitizer, DataFlowSanitizer profile - library which is used to collect coverage information BlocksRuntime - target-independent implementation of Apple "Blocks" runtime interfaces Description Source: https://github.com/conda-forge/compiler-rt-feedstock/blob/main/recipe/meta.yaml | Runtime Library, Compiler Support, Low-Level Support, Language Features | Computer Science, Computer & Information Sciences | Library | Documentation, Uses, and more |
| compiler-rt32 | mcc, lcc | N/A | The compiler-rt32 project is a runtime library that provides functionality for compilers in the 32-bit architecture. It includes various runtime components such as sanitizers, builtins, and support libraries for handling memory operations and error detection. | Compiler, Runtime Library, 32-Bit Architecture | Computer Science, Computer & Information Sciences | Runtime Library | Documentation, Uses, and more |
| compiler32 | mcc, lcc | N/A | compiler32 is a lightweight and efficient compiler software designed for compiling source code into executable programs. It supports various programming languages and optimization techniques to enhance the performance of the compiled code. | Compiler, Software Development, Programming | Computer & Information Sciences | Development Tools | Documentation, Uses, and more |
| compleasm | mcc | 1.0 | compleasm (formerly miniBUSCO) is a fast, efficient tool for assessing the completeness of genome assemblies by quantifying the presence of conserved single-copy orthologs, following the same conceptual approach as BUSCO but with a streamlined, faster implementation. It uses the miniprot protein-to-genome aligner together with BUSCO/OrthoDB lineage marker sets and a hidden-Markov-model-based scoring step to categorize markers as complete (single-copy or duplicated), fragmented, or missing, giving a percentage-based completeness summary. Input is a genome assembly FASTA and a chosen lineage database (auto-downloaded or specified), and output is a summary table plus detailed per-marker results and a full gene table. It offers a run mode for the analysis itself and a separate download step for fetching lineage databases, and it is notably faster than BUSCO on large eukaryotic genomes while producing comparable metrics. This is bioconda compleasm 0.2.6. | genome completeness, assembly evaluation, ortholog analysis | Computational genomics, Bioinformatics | Computational Biology Tool | Documentation, Uses, and more |
| comsol | lcc | N/A | COMSOL Multiphysics 6.1, a finite-element-based multiphysics simulation platform for modeling coupled physical phenomena including structural mechanics, heat transfer, fluid flow, electromagnetics, acoustics, and chemical reactions. Widely used across engineering and applied-physics research for design and analysis. | Simulation, Modeling, Physics, Engineering, Optimization | Mechanical Engineering, Engineering & Technology | Simulation Software | Documentation, Uses, and more |
| conda | lcc | N/A | R | package management, environment management, cross-platform, open-source | Software Engineering, Computer Science | Package Manager | Documentation, Uses, and more |
| conda-bioconductor | mcc | N/A | R | bioinformatics, genomics, data analysis, conda | Bioinformatics, Biological Sciences | Package Manager | Documentation, Uses, and more |
| conmon | ecc | N/A | Conmon is a lightweight tool that provides a way to monitor and manage container processes. It is primarily used in conjunction with container runtimes to provide logging and lifecycle management for containers. | Containerization, Monitoring, Logging, Lightweight | Software Engineering, Other Computer and Information Sciences | Utility | Documentation, Uses, and more |
| coreutils | mcc | N/A | Coreutils is a package that provides the basic file, shell, and text manipulation utilities of the GNU operating system. | GNU, Command Line, Utilities, Open Source | Computer Science, Software Engineering | Command Line Tool | Documentation, Uses, and more |
| cortex_con | mcc | 1.0 | cortex_con is a de novo genome assembler that builds assemblies using colored de Bruijn graphs, a data structure in which k-mers from one or more samples are stored together and each is annotated with which sample(s) (colors) it appears in. It is the 'consensus' variant of the Cortex assembler; while the full Cortex/cortex_var toolkit is designed for variant discovery by comparing colors, cortex_con focuses on assembling a consensus sequence from the graph. It takes sequencing reads, builds a de Bruijn graph at a chosen k-mer size, applies error-cleaning steps (removing low-coverage tips and bubbles), and outputs assembled contigs. Because of the colored-graph foundation it can incorporate multiple read sets, and it was historically used for microbial assembly and as a component in structural-variant and SNP-calling pipelines. This build is version 0.04c from Bioconda; being an older assembler it is typically used for legacy or specialized colored-de-Bruijn-graph workflows rather than as a first-choice modern long-read assembler. | Documentation, Uses, and more | |||
| cortex_var | mcc | N/A | cortex_var software. | Documentation, Uses, and more | |||
| cortexassembler | lcc | N/A | assembler software. | genome assembly, bioinformatics, next-generation sequencing, de novo assembly | Bioinformatics, Biological Sciences | Command-line tool | Documentation, Uses, and more |
| coseg | mcc | 1.0 | COSEG is a program within the RepeatMasker/Dfam transposable-element toolkit that identifies statistically significant subfamilies of a transposable element based on patterns of co-segregating mutations. Starting from a multiple alignment of many individual copies of a repeat family, it detects positions whose mutations tend to occur together across copies—evidence that those copies descend from a distinct ancestral subfamily—and reports the derived subfamily consensus sequences and their relationships. This helps refine TE libraries by splitting a broad family into biologically meaningful lineages. Here it is provided as part of the Dfam TE Tools (TETools) v1.9.5 curated environment packaged from the upstream dfam/tetools:1.95 Docker image, which also contains RepeatMasker, RepeatModeler, and their dependencies. COSEG is typically used by TE-annotation specialists during manual curation of repeat libraries rather than in routine genome-annotation runs. | co-segregation analysis, genomics, genetic variation, population genetics, comparative genomics, bioinformatics | Population Genetics, Bioinformatics | Bioinformatics analysis tool | Documentation, Uses, and more |
| coverm | mcc | 1.0 | CoverM (version 0.7.0 from bioconda) computes read coverage and relative abundance for genomes and contigs from sequencing data, aimed especially at metagenomics where quantifying how much of a community maps to each reference genome or MAG is essential. It can either map reads on the fly (wrapping minimap2 or BWA) or ingest existing BAMs, and offers two main modes: a genome mode for per-genome abundance across a set of reference genomes/bins, and a contig mode for per-contig coverage. It provides many coverage metrics (mean, trimmed mean, relative abundance, RPKM/TPM, covered fraction) selectable by the user, plus read-identity and covered-fraction filters to exclude spurious mappings. It is widely used to profile MAG abundance across metagenomic samples and to prepare coverage inputs for binning. | Metagenomics, Bioinformatics, Genomics | Bioinformatics, Biological Sciences | Genomic Analysis Tool | Documentation, Uses, and more |
| cp2k | mcc, lcc | 1.0 | CP2K is a versatile open-source package for atomistic simulations in quantum chemistry and condensed-matter/solid-state physics, targeting molecular, liquid, and periodic/bulk systems. Its flagship method is the mixed Gaussian and plane-waves (GPW/GAPW) implementation of Kohn-Sham DFT, and it also supports Hartree-Fock, MP2/RPA, semi-empirical and tight-binding methods, and classical force fields, enabling QM/MM and Born-Oppenheimer/ab initio molecular dynamics, geometry optimization, and spectroscopic property calculations. Input is a structured CP2K input file referencing basis sets and pseudopotentials, and calculations are launched in parallel through its parallel psmp executable. This build is CP2K 2026.1 compiled with Spack against GCC 12.5.0 with MPI+OpenMP, libint (for Hartree-Fock exchange) and libxc (exchange-correlation functionals), on an OpenMPI+UCX/InfiniBand stack for scalable multi-node runs. | Computational Chemistry, Quantum Mechanics, Molecular Dynamics | Physical Sciences | Molecular Simulation | Documentation, Uses, and more |
| cpio | mcc | N/A | cpio is a general file archiver and compression utility that is used to manage archives of files. It can create, extract, and list the contents of archives in various formats. | file management, archiving, compression, Unix, Linux | Computer Science | Command-line utility | Documentation, Uses, and more |
| crackling | mcc, lcc | 1.0 | Crackling is a tool for fast, genome-wide design of CRISPR-Cas9 single-guide RNAs (sgRNAs). It scans a target genome to enumerate candidate guides and then ranks them by combining on-target efficiency predictions from multiple consensus scoring models with off-target risk assessment, using an efficient indexing scheme (Inverted Signature Slice Lists) to make genome-wide off-target searching tractable. In this build it uses Bowtie2 for genome alignment/off-target checking and ViennaRNA for secondary-structure-based filtering of guide candidates, retaining only guides that pass efficiency, specificity, and structural criteria. The input is a target genome/sequence with a configuration file specifying the scoring and filtering parameters; the output is a scored, filtered list of recommended guide RNAs. It is installed from the bmds-lab GitHub repository into a Python 3.6 conda environment (with viennarna and bowtie2) as an editable install. | Documentation, Uses, and more | |||
| crest | mcc | 1.0 | CREST (Conformer-Rotamer Ensemble Sampling Tool) is a program from the Grimme group that automates the exploration of a molecule's conformational and rotamer space using fast semiempirical tight-binding methods, principally GFN-xTB via the bundled xtb backend. Its flagship metadynamics-based iMTD-GC workflow discovers low-energy conformers and rotamers, then screens and ranks them by energy to produce a Boltzmann-weighted conformer ensemble suitable for downstream property prediction or higher-level DFT refinement. Beyond conformer searching it supports protonation/deprotomer and tautomer screening, non-covalent complex sampling, and constrained/metadynamics runs. Input is simply a 3D structure, and output includes the ranked ensemble file (crest_conformers.xyz), the best conformer, and population/energy tables. Version 3.0.2 is a core tool in computational chemistry pipelines for generating physically reasonable conformer sets cheaply before expensive quantum-chemistry calculations. | Documentation, Uses, and more | |||
| crossmap | lcc | N/A | Crossmap is a program for genome coordinates conversion between different assemblies. | Genome Assembly, Genomics, Bioinformatics | Genomics, Biological Sciences | Bioinformatics Tool | Documentation, Uses, and more |
| cuda | mcc, lcc | N/A | Nvidia CUDA toolkit and samples | Parallel Computing, Gpu Programming, High Performance Computing, Software Development | Computer Science, Computer & Information Sciences | Compiler | Documentation, Uses, and more |
| cudnn | lcc | 1.0 | Nvidia cudNN library | Deep Learning, Artificial Intelligence, Machine Learning, Gpu Acceleration, Neural Networks | Computer Science | Deep Learning Accelerator | Documentation, Uses, and more |
| cufflinks | mcc, lcc | 1.0 | Cufflinks software. | RNA-Seq, Transcriptomics, Bioinformatics | Biology, Biological Sciences | Tool | Documentation, Uses, and more |
| curl | mcc, ecc | N/A | cURL is a command line tool and library for transferring data with URLs. Description Source: https://curl.se/ |
Networking, Data Transfer, Command-Line, Scripting | Computer Science, Computer & Information Sciences, Other Computer & Information Sciences | Networking Tool | Documentation, Uses, and more |
| cutadapt | mcc, lcc | 1.0 | Cutadapt is a fast, flexible read-trimming tool that finds and removes adapter sequences, primers, poly-A tails, and other unwanted sequence from high-throughput sequencing reads, a standard first step in RNA-seq, small-RNA/miRNA, amplicon, and DNA-seq preprocessing. It supports 3', 5', and anchored/linked adapters with configurable error rates and overlaps, quality trimming of read ends, minimum/maximum length filtering, and full paired-end handling that keeps mates synchronized. It reads FASTA/FASTQ (gzip/bzip2/xz aware) and writes trimmed reads plus a summary report, with separate options for specifying single-end versus paired-end adapters and paired output. This is version 5.1. | Bioinformatics, Ngs, Sequence Analysis, Genomics | Computational Biology, Biological Sciences | Sequence Analysis Tool | Documentation, Uses, and more |
| dakota | lcc | N/A | Dakota is a software toolkit for optimization, uncertainty quantification, and sensitivity analysis. It provides a flexible framework for integrating various analysis codes and facilitates the exploration of design spaces. | Optimization, Uncertainty Quantification, Sensitivity Analysis, Simulation | Engineering, Mechanical Engineering | Toolkit | Documentation, Uses, and more |
| dal | mcc, lcc | N/A | DAL is a library for collision avoidance in robotics applications. | Robotics, Collision Avoidance | Engineering & Technology | Collision Avoidance Software | Documentation, Uses, and more |
| damageproto | mcc | N/A | Documentation, Uses, and more | ||||
| dartrverse | mcc | 1.0 | dartRverse (from CRAN, together with its component packages dartR.base, dartR.data, dartR.popgen, dartR.spatial, dartR.sim, dartR.captive, and dartR.sexlinked) is a modular ecosystem of R packages for analyzing SNP and presence/absence (SilicoDArT) genotype data, particularly data generated by DArTseq, for population, landscape, and conservation genetics. It centers on the genlight data structure and provides an end-to-end workflow: importing DArT reports, filtering by call rate, reproducibility, read depth, minor-allele frequency and linkage, and then running analyses spanning population structure and differentiation (F-statistics, PCoA), diversity and relatedness, spatial/landscape genetics, simulations, captive-breeding management, and detection of sex-linked markers. Analyses use a consistent gl.* function naming (for reading DArT reports, filtering by call rate, running PCoA, and so on), making pipelines readable and reproducible. This container provides it on Rocky 9.3 with R 4.4 alongside Phrynomics and TESS3r and the required geospatial/compiled dependencies. It is popular with conservation geneticists working on non-model species. | Documentation, Uses, and more | |||
| dask | lcc | 1.0 | Dask is a flexible parallel- and distributed-computing library for Python that scales familiar PyData tools beyond a single core or a single machine. It offers high-level collections—dask.array (a chunked, lazy analogue of NumPy), dask.dataframe (a partitioned analogue of pandas), and dask.bag—as well as a low-level delayed/futures interface for parallelizing arbitrary Python code. Computations build a task graph that is executed lazily by a scheduler (threads, processes, or a distributed cluster) only when the result is explicitly requested, enabling out-of-core processing of datasets larger than RAM. A distributed cluster is launched via a scheduler and workers (often through dask.distributed or dask-jobqueue on HPC/SLURM), with a live diagnostics dashboard for monitoring. This is version 2021.2.0, an older but stable release; it is widely used to accelerate ETL, feature engineering, and array analytics without rewriting existing NumPy/pandas code. | Parallel Computing, Scalable Computing, Analytics, Data Science | Data Analytics, Computer & Information Sciences | Library | Documentation, Uses, and more |
| dbus | mcc | N/A | D-Bus is a message bus system, a simple way for applications to talk to one another. In addition to interprocess communication, D-Bus helps coordinate process lifecycle; it makes it simple and reliable to code a "single instance" application or daemon, and to launch applications and daemons on demand when their services are needed. Description Source: https://www.freedesktop.org/wiki/Software/dbus/ |
Message Bus System, Interprocess Communication, Process Lifecycle Management | Software Engineering, Computer & Information Sciences | Middleware | Documentation, Uses, and more |
| debugger | mcc, lcc, ecc | N/A | Intel® oneAPI Application Debugger (gdb-oneapi) | Debugging, Software Development | Engineering & Technology | Debugger | Documentation, Uses, and more |
| deep-md-cpu | lcc | N/A | DeePMD-kit (CPU build) is a machine-learning package for constructing deep-neural-network interatomic potentials from ab initio data and using them to drive molecular dynamics at near-first-principles accuracy. It integrates with MD engines like LAMMPS and is used in computational chemistry and materials science. Installed as a Conda environment. | Deep Learning,Molecular Dynamics,Computational Chemistry | Chemistry, Computational Science | Library | Documentation, Uses, and more |
| deep-md-gpu | lcc | N/A | DeePMD-kit (GPU build) is a deep-learning framework for constructing machine-learning interatomic potentials, enabling molecular dynamics with near-first-principles accuracy at greatly reduced cost. It trains neural network models on ab initio (DFT) data and integrates with MD engines such as LAMMPS and i-PI. Researchers in computational chemistry, materials science, and physics use it to run large-scale, accurate MD simulations of complex atomic systems on GPUs. | Deep Learning, Molecular Dynamics, GPU Computing, Scientific Computing | Molecular Modeling, Computational Chemistry | Simulation Software | Documentation, Uses, and more |
| deepchem | lcc | 1.0 | DeepChem is an open-source Python library that democratizes deep learning for the life sciences, chemistry, materials science, and drug discovery. It provides a large collection of molecular featurizers (converting SMILES, molecular graphs, or 3D structures into fingerprints, graph representations, or Coulomb matrices), curated benchmark datasets (MoleculeNet), data-loading/splitting utilities, and a unified model API wrapping scikit-learn, TensorFlow/Keras, and PyTorch models — including graph convolutional networks, message-passing networks, and other architectures for property prediction. Typical use is programmatic within Python: loading a dataset, featurizing molecules, instantiating a model such as a graph-convolution model, and fitting and predicting to model properties like solubility, toxicity, or binding affinity. It supports classification and regression on molecular and materials data and integrates transfer learning and model evaluation metrics. This is conda-forge deepchem 2.6.1. | deep learning, cheminformatics, drug discovery | Chemoinformatics, Drug Discovery | Library | Documentation, Uses, and more |
| deepchem-2.5.0-gpu | lcc | 1.0 | DeepChem is an open-source Python library for applying machine learning and deep learning to the molecular sciences, including drug discovery, cheminformatics, materials science, quantum chemistry, and biology. It provides featurizers that turn molecules into ML-ready representations (ECFP/circular fingerprints, graph representations, molecular descriptors), curated dataset loaders and benchmark suites (MoleculeNet), data splitters (including scaffold splitting), and a library of models spanning classical methods and neural architectures such as graph convolutional and message-passing networks for property prediction. This GPU build pairs DeepChem 2.5.0 with TensorFlow 2.4 (GPU-enabled) for accelerated training and RDKit 2020.09 for molecular handling and featurization, exposed through the deepchem Python API (e.g. dc.data, dc.feat, dc.models). | Deep Learning, Bioinformatics, Chemoinformatics, Machine Learning | Bioinformatics, Drug Discovery | Library | Documentation, Uses, and more |
| deeplabcut | lcc | 1.0 | DeepLabCut is a markerless pose-estimation toolbox for quantifying animal behaviour from video, widely used in neuroscience and ethology. It tracks user-defined body parts frame by frame using deep neural networks, turning raw video into coordinate time series suitable for kinematic and behavioural analysis. The standard workflow creates a project, extracts and labels a small set of representative frames, trains a network on a GPU, evaluates accuracy against held-out labels, and then analyses whole videos to produce per-frame body-part coordinates with confidence values. Version 3 runs on a PyTorch backend and also provides SuperAnimal models, pre-trained networks for rodents, quadrupeds, birds and humans that can annotate video with no manual labelling at all; the rodent and quadruped SuperAnimal weights are included in this deployment so analyses run without downloading anything. Outputs are stored as HDF5 and CSV tables of body-part positions, optionally with labelled videos and trajectory plots. | Documentation, Uses, and more | |||
| deeppolisher | mcc | 1.0 | DeepPolisher is a deep-learning genome-assembly polisher from Google's Genomics team that corrects base-level errors (mismatches and small indels) in long-read assemblies by learning to predict the correct sequence from patterns in read alignments. It uses a transformer-based model that examines pileups of reads (typically PacBio HiFi, produced by aligning reads back to the draft assembly with pbmm2/minimap2) and emits corrected consensus sequence, improving per-base quality (QV) of haplotype-resolved assemblies such as those from hifiasm or Verkko. The workflow aligns reads to the assembly, runs DeepPolisher's image-generation and inference steps, and applies the predicted edits to produce a polished FASTA, with a companion step for phasing-aware polishing. This build is the CPU-only installation. It is used in high-quality reference and pangenome assembly pipelines, including those associated with the Human Pangenome Reference Consortium, to squeeze the last errors out of near-complete assemblies. | genome polishing, long-read sequencing, deep learning, neural networks, genome assembly, error correction, bioinformatics | Genome Assembly, Bioinformatics | Analysis software | Documentation, Uses, and more |
| deeptools | mcc, lcc | 1.0 | deepTools is a suite of command-line tools for processing and visualizing high-throughput sequencing coverage data, especially from ChIP-seq, ATAC-seq, RNA-seq, and related assays. It converts and normalizes aligned reads into coverage tracks with bamCoverage (supporting RPKM/CPM/BPM/RPGC normalization, extension, and bin sizes), compares two samples (e.g. treatment vs. input) with bamCompare, and assesses experiment quality with tools like multiBamSummary, plotCorrelation, plotPCA, plotFingerprint, and plotCoverage. Its signature visualization workflow uses computeMatrix to summarize one or more bigWig signals over sets of genomic regions (e.g. gene bodies or TSSs from a BED/GTF) and then plotHeatmap or plotProfile to render metaprofiles and heatmaps of signal enrichment. Inputs are indexed BAM files and region files; outputs are bigWig tracks, matrices, and publication-ready plots. Version 3.4.3 is provided as its own conda environment exposed as an %apprun within a multi-app Miniconda container that also bundles climate-science tools. | Bioinformatics, Hpc Tools, Computational Software | Biology, Biological Sciences | Analysis Tool | Documentation, Uses, and more |
| deepvariant | mcc, lcc | 1.0 | DeepVariant is a deep-learning-based small-variant caller from Google that reframes variant calling as an image-classification problem: it encodes candidate variant loci from aligned reads as pileup image tensors and uses a trained convolutional neural network to classify each as homozygous reference, heterozygous, or homozygous alternate, achieving high accuracy for SNPs and indels. It supports data from multiple sequencing technologies through model types — Illumina whole-genome (WGS) and exome (WES), PacBio HiFi, and a hybrid PacBio-Illumina mode in this 1.2.0 release (dedicated Oxford Nanopore model support was added in later versions) — and runs a three-stage pipeline (make_examples, call_variants, postprocess_variants) that takes a reference FASTA and a sorted, indexed BAM/CRAM and produces a standard VCF (and optionally gVCF). It is typically run via the run_deepvariant wrapper, selecting the model type appropriate to the data. This container is built from the official Google DeepVariant 1.2.0 image with added CCS cluster mount points. | Variant Calling, Genetic Variants, Deep Learning, DNA Sequencing | Genomics, Biological Sciences | Variant Caller | Documentation, Uses, and more |
| delly | lcc | 1.0 | DELLY is a structural-variant (SV) discovery and genotyping tool for whole-genome resequencing data, operating in the domain of population and cancer genomics. It integrates paired-end read-pair signals and split-read alignments to detect deletions, tandem duplications, inversions, and both intra- and inter-chromosomal translocations at single-nucleotide breakpoint resolution, and it can additionally call copy-number variants. It takes coordinate-sorted, indexed BAM/CRAM files plus the indexed reference FASTA as input and writes SV calls to BCF/VCF. A typical workflow first calls SVs per sample, optionally merges sites across a cohort and re-genotypes them, and then filters somatic or germline variants with its filter step. It is commonly paired with bcftools for downstream conversion and filtering. This is version 1.3.1. | Structural Variant Discovery, Sv Detection, Genomics | Genomics, Biological Sciences | Tool | Documentation, Uses, and more |
| desktop | mcc, ecc | 1.0 | Desktop is a browser-based interactive Open OnDemand application (the standard bc_desktop app) that launches a full graphical Linux desktop session running on an MCC compute node and streams it to the user's web browser via OnDemand's noVNC gateway. It lets researchers use GUI software — file managers, graphical editors, molecular viewers, IDEs, and other X11 applications — on allocated HPC resources without configuring their own VNC or X forwarding. The user requests a session through the OnDemand portal (mcc-ood.ccs.uky.edu), choosing a partition/account, number of cores, memory, and wall time, and OnDemand submits the underlying SLURM job and provides a Launch button once the node is ready. In the SDS context this entry is a catalog-only stub: it is not built into a container image but is registered so the interactive desktop appears in the Software Discovery Service alongside command-line tools, improving discoverability of MCC's OnDemand offerings. | Documentation, Uses, and more | |||
| detonate | lcc | 1.0 | DETONATE (DE novo TranscriptOme rNa-seq Assembly with or without the Truth Evaluation), version 1.11, provides reference-free, model-based evaluation of de novo transcriptome (and other de novo) assemblies. Its two components are RSEM-EVAL, which computes a single principled score for an assembly by modeling how well it explains the RNA-seq reads (jointly accounting for read alignment, contig length, and assembly compactness), and REF-EVAL, a suite of alignment-based measures for comparing an assembly against a reference or another assembly. RSEM-EVAL is especially useful because it lets you rank alternative assemblies, tune assembler parameters, or choose the best assembly without needing a reference genome. Inputs are the assembled contigs FASTA and the RNA-seq reads; the tool outputs a numeric score and per-contig contributions. It is commonly used in non-model-organism transcriptomics to benchmark Trinity or other de novo assemblies. | Documentation, Uses, and more | |||
| dev-utilities | mcc, lcc, ecc | N/A | Sample headers and CLI sample browser (oneapi-cli). | Development, Utilities | Computer & Information Sciences, Other Computer & Information Sciences | Tools | Documentation, Uses, and more |
| diamond | mcc, lcc | 1.0 | diamond software. | Sequence Alignment, Ngs Analysis, Bioinformatics | Bioinformatics, Biological Sciences | Bioinformatics Tool | Documentation, Uses, and more |
| diffdock | lcc | 1.0 | DiffDock is a state-of-the-art deep-learning method for blind molecular docking, predicting how a small-molecule ligand binds to a protein without requiring a known binding pocket. It frames docking as a diffusion generative process over the space of ligand poses (translation, rotation, and torsion angles), sampling multiple candidate poses and then ranking them with a learned confidence model, which yields a confidence score for each prediction. Inputs are a protein structure (PDB) or sequence and a ligand as SMILES or SDF; for sequence input it folds the protein on the fly using ESMFold/OpenFold. It is run through its inference script over a CSV of protein-ligand pairs and outputs ranked ligand poses as SDF files. This GPU build is release v1.1.3 with the v1.1 pretrained weights, using PyTorch 1.13.1+cu117 and PyTorch Geometric. | Molecular Docking, Bioinformatics, Computational Chemistry | Bioinformatics, Molecular Docking | Application | Documentation, Uses, and more |
| diffutils | mcc, lcc, ecc | N/A | GNU Diffutils is a package of several programs related to finding differences between files. Description Source: https://github.com/conda-forge/diffutils-feedstock/blob/main/recipe/meta.yaml | Text Comparison, File Difference, Text Analysis, Version Control | Software Engineering, Engineering & Technology | Text Processing | Documentation, Uses, and more |
| dimemas | mcc, lcc | N/A | Dimemas tool | Performance Prediction, Parallel Applications, Simulation, Large-Scale Systems | Computer Science, Engineering & Technology | Performance Prediction Tool | Documentation, Uses, and more |
| direnv | mcc, lcc | N/A | direnv 2.37.1 - per-directory environment switcher for the shell. It loads and unloads environment variables, PATH entries and environment modules from a project's .envrc file as you enter and leave that directory, so each project can carry its own environment without editing your shell startup files. Also usable non-interactively in batch jobs via 'direnv exec'. | Documentation, Uses, and more | |||
| dispy | lcc | N/A | dispy is a generic, comprehensive, yet easy to use framework and tools for creating, using and managing compute clusters to execute computations in parallel across multiple processors in a single machine (SMP), among many machines in a cluster, grid or cloud. | distributed computing, parallel processing, Python, task scheduling, cluster computing, HPC | High Performance Computing, Computer Science | Library | Documentation, Uses, and more |
| dnnl | mcc, lcc, ecc | N/A | Performance library of basic building blocks for deep learning applications | Deep Learning, Neural Networks, Machine Learning, Performance Optimization | Artificial Intelligence & Intelligent Systems, Computer & Information Sciences | Deep Learning Library | Documentation, Uses, and more |
| dnnl-cpu-gomp | mcc, lcc | N/A | The dnnl-cpu-gomp is an open-source deep neural network library developed for CPU-based computations employing OpenMP (Open Multi-Processing) as a thread-processing technology. | Deep Learning, Neural Networks, Openmp, Cpu Computing | Machine Learning, Computer & Information Sciences | Computational Software | Documentation, Uses, and more |
| dnnl-cpu-iomp | mcc, lcc | N/A | dnnl-cpu-iomp is an optimization library for deep neural network computations on CPU using Intel OpenMP (iomp) for improved performance. | Deep Learning, Optimization, Cpu Acceleration | Computer & Information Sciences | Documentation, Uses, and more | |
| dnnl-cpu-tbb | mcc, lcc | N/A | DNNL (Deep Neural Network Library) with TBB (Threading Building Blocks) support for optimized performance on CPU architectures. | Deep Learning, High-Performance Computing, Multi-threading, CPU Optimization | Artificial Intelligence, Artificial Intelligence and Intelligent Systems | Library | Documentation, Uses, and more |
| docbook-xml | mcc | N/A | DocBook XML is a semantic markup language for technical documentation, providing a framework for writing structured documents. | XML, Documentation, Technical Writing | Computer Science | Markup Language | Documentation, Uses, and more |
| docbook-xsl | mcc | N/A | DocBook XSL is a set of stylesheets for transforming DocBook XML documents into various output formats, such as HTML, PDF, and more. | XML, XSLT, Documentation, Publishing | Computer Science, Computer and Information Sciences | Documentation Tool | Documentation, Uses, and more |
| dock6 | lcc | N/A | DOCK 6.9 is a molecular docking application used in computer-aided drug design. It predicts the binding orientation and affinity of small-molecule ligands to receptor targets, supporting virtual screening, pose generation, and scoring in computational chemistry and pharmacology. | molecular docking, virtual screening, drug discovery, protein–ligand interactions, computational chemistry, structure-based design | Chemistry, Computational Chemistry | Scientific Software | Documentation, Uses, and more |
| dorado | mcc, lcc | 1.0 | Dorado is Oxford Nanopore Technologies' current-generation, open-source basecaller that converts the raw electrical signal (POD5/FAST5) produced by nanopore sequencers into nucleotide sequences. It is GPU-accelerated (CUDA) and ships with the high-accuracy neural-network models (fast, hac, and sup) and supports modified-base calling such as 5mC/5hmC and 6mA. Beyond basecalling, dorado provides subcommands for duplex calling (higher-accuracy reads from complementary strands), demultiplexing barcoded runs, adapter/primer trimming, poly(A) estimation, and even direct alignment via an embedded minimap2. Its basecaller emits an unaligned BAM with quality scores and optional modified-base tags. Version 0.9.5 is delivered as a prebuilt Linux binary and is the recommended replacement for the older Guppy/Bonito basecallers in long-read genomics pipelines. | Performance Evaluation, Benchmarking, Hpc, Supercomputing | Biology, Engineering & Technology | Tool | Documentation, Uses, and more |
| dorodo | mcc | 1.0 | Dorado is Oxford Nanopore Technologies' high-performance, GPU-accelerated basecaller that converts the raw electrical signal (POD5/FAST5) produced by nanopore sequencing into nucleotide sequences. It supports simplex basecalling, duplex basecalling (which pairs template and complement strands for substantially higher accuracy), and modified-base calling (e.g. 5mC/5hmC methylation) using downloadable neural-network models of varying speed/accuracy (fast, hac, sup). It can also demultiplex barcodes, trim adapters, and align to a reference in-line, emitting unaligned or aligned BAM (or FASTQ). A run first downloads a model, then performs simplex basecalling over a POD5 directory (optionally enabling modified-base calling for methylation), while a separate duplex mode handles paired-strand basecalling. This deployment stages the version 0.3.2 Linux x64 release tarball; note the catalog name 'dorodo' is a typo for 'dorado'. It is provided within a multi-app Miniconda container on Rocky 8 that bundles long-read (Nanopore/PacBio) and methylation tools, each as its own environment/app. | Documentation, Uses, and more | |||
| double-conversion | lcc | N/A | Double-Conversion provides binary-decimal and decimal-binary routines for IEEE doubles. This library consists of efficient conversion routines that have been extracted from the V8 JavaScript engine. Description Source: https://github.com/google/double-conversion |
Conversion, Performance, Precision | Numerical Analysis, Computer & Information Sciences | Library | Documentation, Uses, and more |
| download_eggnog_data | mcc | 1.0 | download_eggnog_data.py is the database-staging helper script that ships with eggNOG-mapper (bioconda eggnog-mapper 2.1.15). eggNOG-mapper is a fast tool for functional annotation of large sets of protein or CDS sequences based on precomputed orthology assignments from the eggNOG database of orthologous groups; before it can run, its reference data must be downloaded locally, which is exactly what this script does. It fetches and installs the eggNOG annotation database, the eggNOG proteins/taxa data, and, depending on options, the Diamond or MMseqs2 search databases and optional extras such as the HMM databases, Pfam data, or the MMseqs database. Command-line flags control which components are retrieved and where they are stored — for example a non-interactive mode with a chosen data directory downloads the core data, with additional options to add Pfam, MMseqs, or per-taxon HMM data. This step is a prerequisite for running the emapper annotation tool, which then annotates query sequences with orthologous groups, GO terms, KEGG pathways/KOs, EC numbers, and COG categories. | Documentation, Uses, and more | |||
| dpct | mcc, lcc, ecc | N/A | Migrate existing CUDA* code to SYCL code. | Cuda Programming, Gpu Computing, Parallel Programming | Computer Science, Computer & Information Sciences | Development Compiler | Documentation, Uses, and more |
| dpl | mcc, lcc, ecc | N/A | Intel(R) oneAPI DPC++ Library provides an alternative for C++ developers who create heterogeneous applications and solutions. Its APIs are based on familiar standards - C++ STL, Parallel STL (PSTL), Boost.Compute, and SYCL* - to maximize productivity and performance across CPUs, GPUs, and FPGAs. | Programming Language Analysis, Syntax Analysis, Language Structure Analysis | Documentation, Uses, and more | ||
| dram | mcc | 1.0 | DRAM (Distilled and Refined Annotation of Metabolism), version 1.5.0, is a tool for annotating the functional and metabolic potential of microbial genomes, isolate genomes, viral contigs, and especially metagenome-assembled genomes (MAGs). It annotates predicted genes against multiple databases — KEGG (KOfam), Pfam, dbCAN (CAZymes), MEROPS peptidases, UniRef, VOGDB, and others — and then "distills" the raw annotations into interpretable, curated summaries of metabolic pathways and functions, producing per-genome and cross-genome product tables plus an interactive heatmap of metabolic modules. A companion, DRAM-v, is tailored to viral sequences (from VirSorter output) and identifies auxiliary metabolic genes. The typical workflow is a two-stage process in which an annotate step assigns database annotations to predicted genes and a distill step condenses those annotations into curated metabolic summaries. Note that DRAM depends on large, separately downloaded reference databases. It is heavily used in environmental and host-associated microbiome studies to compare metabolic capabilities across genomes. | Documentation, Uses, and more | |||
| drep | mcc | 1.0 | dRep, version 3.5.0, rapidly compares large collections of genomes and dereplicates them — identifying groups of essentially the same genome and selecting a single high-quality representative from each — which is essential when many metagenome-assembled genomes (MAGs) or isolates redundantly represent the same organisms. It uses a fast two-stage clustering strategy: a quick primary clustering with Mash to group broadly similar genomes, followed by a precise secondary clustering with a genome-alignment ANI method (e.g. fastANI or gANI) at a user-set threshold (commonly 99% for strains, 95% for species). It integrates CheckM completeness/contamination scores into a composite quality score to choose the best representative per cluster. Its dereplicate mode compares genomes and chooses representatives, while a separate compare mode performs clustering only. Outputs include dereplicated genome sets, cluster assignments, and diagnostic plots, making it a standard step in metagenomic MAG curation. | Microbiome, Genome Analysis, Bioinformatics | Bioinformatics, Biological Sciences | Bioinformatics | Documentation, Uses, and more |
| dsuite | mcc, lcc | 1.0 | Dsuite is a fast C++ program for detecting and quantifying introgression (gene flow/admixture) between populations or species directly from population-genomic data. Operating on a multi-sample VCF plus a file assigning samples to populations, it computes Patterson's D-statistic (the ABBA-BABA test) and the related f4-ratio and f-branch (fb) statistics for all trios (or a user-supplied tree topology) in a single pass, avoiding the need for expensive per-trio resampling tools. Its main subcommands are Dtrios (compute D and f4-ratio across all population trios, with block-jackknife standard errors and Z-scores), Fbranch (assign admixture signals to specific branches of a phylogeny), and Dinvestigate (sliding-window fdM/fd/df scans to localize introgressed genomic regions), producing tables of D, f4-ratio, p-values and BBAA/ABBA/BABA counts. It is widely used in speciation and phylogeographic studies to test the assumption of a strictly bifurcating tree. | Python Library, Data Analysis, Data Visualization | Biostatistics, Statistics & Probability | Python Library | Documentation, Uses, and more |
| e2fsprogs | ecc | N/A | e2fsprogs is a set of utilities for creating, checking, and maintaining the ext2, ext3, and ext4 file systems in Linux. | Linux, File System, Utilities | Operating Systems, File Systems | Utility | Documentation, Uses, and more |
| eaglev | mcc | 1.0 | Eagle, version 2.4.1 (Po-Ru Loh, Broad Institute / Alkes Price group), performs fast and accurate statistical phasing of genotype data — inferring which alleles at heterozygous sites lie on the same parental haplotype. It offers two modes: reference-based phasing, which phases a target sample set against a large external haplotype reference panel (e.g. HRC or 1000 Genomes) using a fast HMM, and reference-free phasing, which leverages long identity-by-descent sharing within very large cohorts to phase without any reference. Its Positional Burrows-Wheeler Transform-based algorithms make it scalable to biobank-sized datasets. Input is genotypes in VCF/BCF (or PLINK bed for reference-free mode) plus a genetic map; output is a phased VCF with haplotype-resolved genotypes. Accurate phasing from Eagle is a standard precursor to genotype imputation (e.g. with Minimac/IMPUTE) and haplotype-based analyses. | Documentation, Uses, and more | |||
| earlgrey | mcc | 1.0 | Earl Grey is an automated, end-to-end transposable-element (TE) annotation pipeline that streamlines the discovery, curation, and quantification of repetitive elements in genome assemblies while addressing common problems such as fragmented and misannotated repeats. It runs de novo repeat identification (RepeatModeler), performs an automated 'BLAST, Extract, Extend' process to recover full-length TE consensus sequences, then annotates and defragments repeats (RepeatMasker/RepeatCraft-style merging) and produces summary tables, repeat landscape plots, and GFF outputs. It takes a genome FASTA plus a species/prefix argument and runs as a single command that orchestrates the whole workflow. It is provided here from the prebuilt wilsonleunggep/earlgrey:v0 Docker image, bundling its many dependencies together. | General Tags transposable elements, repeat annotation, genome annotation, de novo TE discovery, eukaryotic genomes, bioinformatics pipeline | Evolutionary Genomics, Bioinformatics | Analysis software | Documentation, Uses, and more |
| easybuild | mcc, lcc, ecc | N/A | Software build and installation framework | Software Building, High-Performance Computing, Automation | Computer Science, Engineering & Technology | Build Automation | Documentation, Uses, and more |
| ecc/cuda | ecc | N/A | NVIDIA CUDA Toolkit 12.8.0 installed at /share/apps/cuda/12.8.0_570.86.10 | Documentation, Uses, and more | |||
| ecc/miniconda3 | ecc | N/A | Miniconda is a minimal installer for conda, a package manager and environment manager. | Documentation, Uses, and more | |||
| ecc/ollama | ecc | N/A | rclone is a command-line tool for launching LLMs. | Documentation, Uses, and more | |||
| ecc/rclone | ecc | N/A | Rclone is a command-line tool for syncing files to and from cloud storage. | Documentation, Uses, and more | |||
| ecc/spack | ecc | N/A | Spack 0.23.1 is a flexible package manager for HPC environments. | Documentation, Uses, and more | |||
| ed | ecc | N/A | ed is a line-oriented text editor for Unix systems, designed for editing text files in a simple and efficient manner. | Text Editor, Unix, Command Line | Software Engineering, Other Computer and Information Sciences | Text Editing Tool | Documentation, Uses, and more |
| eggnog-mapper | mcc | 1.0 | eggNOG-mapper performs fast, large-scale functional annotation of protein and nucleotide sequences using precomputed orthology assignments from the eggNOG database of orthologous groups. Instead of costly de novo inference, it maps query sequences to eggNOG orthologs (via DIAMOND, MMseqs2, or HMMER search) and transfers functional information — Gene Ontology terms, KEGG pathways/KOs, COG functional categories, EC numbers, CAZy, Pfam domains, and free-text descriptions — from the finest resolvable orthologous group, which improves specificity over simple best-hit transfer. It is widely used to annotate genomes, metagenomes, and transcriptomes. The main tool, emapper.py, takes a FASTA of proteins (or genes/contigs with built-in gene prediction) and writes tab-delimited annotation tables plus optional GFF/orthology outputs; the eggNOG reference data must be downloaded once with a companion download utility. This is version 2.1.12 from bioconda. | Functional Annotation, Orthology, Gene Function Prediction | Bioinformatics, Biological Sciences | Bioinformatics Tool | Documentation, Uses, and more |
| eigen | mcc, lcc | N/A | eigen free peer-reviewed portable C++ source libraries | Linear Algebra, Numerical Computation, C++ Library | Applied Mathematics, Computer & Information Sciences | Linear Algebra Library | Documentation, Uses, and more |
| elfutils | ecc | N/A | elfutils is a collection of utilities and libraries to read, create and modify ELF binary files, find and handle DWARF debug data, symbols, thread state and stacktraces for processes and core files on GNU/Linux. Description Source: https://sourceware.org/elfutils/ |
Binary Tool, Elf Files, Debugging, Programming | Software Engineering, Other Computer and Information Sciences | Utility | Documentation, Uses, and more |
| elpa | mcc | N/A | The publicly available ELPA library provides highly efficient and highly scalable direct eigensolvers for symmetric (hermitian) matrices. Descripiton Source:https://elpa.mpcdf.mpg.de/ABOUT_ELPA.html | Linear Algebra, Eigenvalue Problems, High-Performance Computing, Scalable Algorithms | Numerical Linear Algebra, Computer & Information Sciences | Library | Documentation, Uses, and more |
| em | lcc | N/A | ANSYS EM (Electromagnetics) version 21.2, part of the ANSYS Electronics/electromagnetic simulation suite. Used by engineers for high-frequency and low-frequency electromagnetic field analysis, antenna, RF, microwave, and electronic component design. A real user-facing engineering simulation application. | Documentation, Uses, and more | |||
| emboss | lcc | N/A | R is a language and environment for statistical computing and graphics (S-Plus like). | Bioinformatics, Computational Biology, Sequence Analysis, Biological Data | Bioinformatics, Biological Sciences | Analysis Tool | Documentation, Uses, and more |
| ensembl-vep | mcc, lcc | 1.0 | Ensembl VEP, the Variant Effect Predictor (version 104.3 from bioconda), annotates genomic variants and predicts their molecular consequences on genes, transcripts, and proteins. Given variants in VCF (or other formats), it reports affected transcripts and consequence terms (e.g., missense, stop-gained, splice-site), amino-acid changes, and can add rich annotations including SIFT/PolyPhen pathogenicity scores, allele frequencies from population databases, regulatory-region overlaps, and known variant identifiers, extensible through a plugin system and custom annotation files. It can run using a locally installed cache, downloaded offline, or query Ensembl databases, and supports GRCh37/GRCh38 and many species. It is one of the most widely used variant-annotation engines in clinical and research genomics, and integrates with tools like bcftools and downstream filtering pipelines. | bioinformatics, genomics, variant analysis | Biology, Genetics | Command-line tool | Documentation, Uses, and more |
| ensembletr | mcc | 1.0 | EnsembleTR (from PyPI, installed in a conda environment alongside htslib, pgenlib, and cyvcf2) merges and harmonizes tandem-repeat (TR/STR) genotype calls produced by multiple independent TR genotyping tools into a single consensus call set. Different STR callers (such as HipSTR, GangSTR, ExpansionHunter, and adVNTR) use differing coordinate and allele conventions, and EnsembleTR reconciles overlapping loci, normalizes allele representations, and applies an ensemble voting/consensus procedure to increase genotype accuracy and consistency across a cohort. It takes as input the per-caller VCFs (and a reference) and outputs a unified, standardized VCF of consensus TR genotypes suitable for downstream population and association analyses. It was developed in the context of large-scale human TR studies where combining callers improves reliability at these historically difficult-to-genotype loci. The htslib/cyvcf2/pgenlib dependencies provide fast VCF/PLINK I/O for large sample sets. | transcriptomics, RNA-seq, ensemble methods, gene expression, transcript reconstruction, bioinformatics | RNA Sequencing Analysis, Bioinformatics | Analysis software | Documentation, Uses, and more |
| entap | mcc | 1.0 | EnTAP (Eukaryotic Non-Model Transcriptome Annotation Pipeline) is a functional-annotation pipeline optimized for de novo assembled transcriptomes of non-model eukaryotes, where a reference genome and curated gene models are unavailable. It streamlines annotation by first filtering the assembly to remove likely non-expressed or fragmented transcripts and to select the best representative isoforms, optionally screening out contaminant sequences, then performing similarity search against protein databases (via DIAMOND) and integrating orthology/gene-family and protein-domain information (e.g. EggNOG, InterPro) to assign gene names, GO terms, and pathway annotations, producing a consolidated, filtered set of annotated coding transcripts. Inputs are a transcriptome FASTA (and optional predicted proteins and expression/abundance data); outputs are annotation tables and filtered FASTA files. It runs in a protein or nucleotide annotation mode after a configuration step that formats the databases. This deployment wraps the plantgenomics/entap Docker image as a Singularity container; note that the def's %labels Name is incorrectly set to 'braker', which does not reflect the actual EnTAP contents. | transcriptome annotation, non-model organisms, functional annotation, ortholog identification, RNA-seq, bioinformatics pipeline | Genome Automation, Bioinformatics | Analysis software | Documentation, Uses, and more |
| esm | lcc | 1.0 | This is the BioHub ESM engine, a protein language-model toolkit built on the Evolutionary Scale Modeling family of models for computational protein analysis on GPUs. It provides ESMFold-style 3D structure prediction (labeled ESMFold2 here), which predicts a protein's atomic structure directly from a single amino-acid sequence without requiring a multiple-sequence alignment, and ESMC (ESM Cascade) per-residue embeddings, which encode each residue into high-dimensional vectors that capture structural and functional context for downstream tasks such as variant-effect prediction, function annotation, remote-homology detection, and as features for machine-learning models. Inputs are protein sequences (FASTA); outputs are predicted structures (e.g. PDB) and/or numerical embedding tensors, and the engine is oriented toward batch modeling of many proteins. It is installed from a pinned BioHub fork of ESM at a fixed commit alongside a pinned BioHub transformers fork and torch 2.6.0+cu124, and is deployed in a CUDA 12.4 GPU container for batch protein modeling on the LCC cluster. | Earth System Modeling, Climate Simulation, Environmental Processes, Earth System Interactions | Earth & Environmental Sciences | Simulation Tool | Documentation, Uses, and more |
| exabayes | mcc, lcc | N/A | ExaBayes is a software package used for Bayesian inference of phylogenetic trees and parameters. | Bayesian Inference, Phylogenetics, Parallel Computing, Mcmc, Hpc | Ecology, Biological Sciences | Phylogenetic Software | Documentation, Uses, and more |
| excerpt | mcc | 1.0 | exceRpt (extra-cellular RNA processing toolkit) is an end-to-end pipeline for processing and quantifying small-RNA and extracellular-RNA (exRNA) sequencing libraries, developed as part of the NIH Extracellular RNA Communication Consortium and widely used through the exRNA Atlas. Starting from raw FASTQ, it performs adapter/3' clipping and QC, filters common contaminants (ribosomal, UniVec), and then hierarchically aligns reads against a series of endogenous references to quantify multiple RNA biotypes — miRNAs (miRBase), tRNAs, piRNAs, other small RNAs, and mRNA/lincRNA — followed by mapping of remaining reads to exogenous genomes/ribosomal sequences for putative microbial and dietary content. It outputs per-sample read counts and biotype breakdowns plus rich diagnostic QC, and ships a companion R merge/processing step to combine many samples into expression matrices. Here it is delivered by wrapping the upstream rkitchen/exceRpt Docker image; a run mounts the input FASTQ and a prebuilt reference database directory and specifies adapter and genome (e.g. hg38) parameters. | text processing, field extraction, command-line, Unix utilities, shell pipelines, data wrangling | Software Utilities, Computer Science | Command-line utility | Documentation, Uses, and more |
| exonerate | lcc | 1.0 | Genomic software . | Sequence Alignment, Homology Search, Bioinformatics, Genomics | Bioinformatics, Biological Sciences | Bioinformatics Tool | Documentation, Uses, and more |
| expat | mcc, lcc, ecc | N/A | Expat is a stream-oriented XML parser library written in C and excels with files too large to fit RAM, and where performance and flexibility are crucial. Description Source: https://libexpat.github.io/ |
Xml Parser, Library | Software Engineering, Computer & Information Sciences | Xml Parser | Documentation, Uses, and more |
| extrae | mcc, lcc | N/A | Extrae tool | Performance Analysis, Trace Data, Parallel Applications, Distributed Applications | Computer Science, Computer & Information Sciences, Software Engineering, Systems & Development, Engineering & Technology, Training, Infrastructure & Instrumentation | Tool | Documentation, Uses, and more |
| fast5 | lcc | 1.0 | fast5 is a lightweight C++/Python library and set of command-line tools for reading and working with the Oxford Nanopore FAST5 file format, which stores raw and processed nanopore signal data in HDF5 containers. It provides programmatic access to the raw current ('squiggle') samples, event data, base-called sequences, and the various metadata and analysis groups embedded in FAST5 files, giving developers a convenient API to extract and inspect nanopore data without directly wrestling with the raw HDF5 hierarchy. It is used as a building block in nanopore analysis pipelines and for custom signal-level processing or format conversion. This is version 0.6.5 from bioconda; note that ONT has largely transitioned to the newer POD5 format, but FAST5 remains common in existing datasets. | Documentation, Uses, and more | |||
| fastani | lcc | 1.0 | FastANI, version 1.34, computes whole-genome Average Nucleotide Identity (ANI) between prokaryotic genomes using a fast, alignment-free approach based on Mashmap-style approximate mapping of genome fragments. ANI is the leading operational metric for microbial species delineation — a ~95% ANI threshold corresponds well to species boundaries — and FastANI makes it feasible to compute across thousands of genomes because it avoids full genome alignment. It accepts assembled genomes as FASTA and supports one-to-one, one-to-many (via query/reference lists), and many-to-many all-vs-all comparisons, reporting for each pair the ANI value and the count/fraction of orthologous fragments used, which indicates alignment confidence. It is widely used for genome dereplication, taxonomic assignment, and building ANI similarity matrices for clustering. | Bioinformatics, Genome Comparison, Bacterial Genomes | Biology, Biological Sciences | Genome Analysis Tool | Documentation, Uses, and more |
| fastp | mcc, lcc | 1.0 | fastp is an ultra-fast, all-in-one FASTQ preprocessing tool written in C++ that combines quality control, adapter trimming, and read filtering into a single multithreaded pass over the data. For single- or paired-end reads it performs automatic adapter detection and trimming (adapters can also be specified explicitly), per-read and sliding-window quality trimming, polyG/polyX tail trimming (important for NovaSeq/NextSeq two-color chemistry), length and quality filtering, base correction in overlapping paired-end regions, and optional UMI processing and deduplication. It produces cleaned FASTQ output plus rich HTML and JSON reports summarizing before/after quality metrics, GC content, duplication, and adapter content — effectively giving FastQC-style QC and trimming in one command. It is a common first step in essentially any short-read (Illumina) sequencing pipeline. This is version 0.23.4. | Bioinformatics, Sequencing, Data Processing | Bioinformatics, Biological Sciences | Tool | Documentation, Uses, and more |
| fastplong | mcc | 1.0 | fastplong is the long-read counterpart to the popular fastp tool from the OpenGene group, providing all-in-one, ultra-fast quality control, adapter trimming, and filtering for Oxford Nanopore and PacBio reads. It performs quality profiling, per-read and per-base quality filtering, length filtering, adapter detection and trimming (including auto-detection for long-read adapters), and generates comprehensive HTML and JSON reports summarizing read-quality distributions before and after processing. Input is one FASTQ file (long reads are single-end) and output is a cleaned FASTQ plus report files, with options controlling minimum length, mean quality, and adapter sequences. It is designed to consolidate several QC steps into a single fast pass over the data. This build is bioconda fastplong 0.4.1, augmented with the parallel.py helper from OpenGene/fastplong v0.4.1. | Documentation, Uses, and more | |||
| fastqc | lcc | 1.0 | FastQC (version 0.11.9 from bioconda) is the standard quality-control tool for high-throughput sequencing data, generating a concise per-file report that flags potential problems before downstream analysis. From FASTQ (or SAM/BAM) input it computes modules including per-base and per-sequence quality scores, per-base sequence content, GC distribution, sequence-length distribution, per-base N content, overrepresented sequences and adapter contamination, and sequence duplication levels, each summarized with pass/warn/fail indicators. It runs either in an interactive GUI or, more commonly on HPC, in batch mode across one or more FASTQ files, producing an HTML report plus a zipped data folder per file. Its outputs are frequently aggregated across many samples with MultiQC. It is almost universally the first step in sequencing pipelines to decide on trimming and to detect library-prep or sequencer issues. | Quality Control, High Throughput Sequencing, Sequence Analysis, Bioinformatics | Biology, Biological Sciences | Analysis Tool | Documentation, Uses, and more |
| fastsimcoal2 | mcc | 1.0 | fastsimcoal2 (v28 Linux 64-bit binary, from the Excoffier lab) is a fast continuous-time coalescent simulator and demographic-inference engine used in population genetics to estimate the parameters of complex evolutionary models. It can simulate genetic data (DNA, SNPs, microsatellites) under arbitrarily complex demographic scenarios — including population size changes, splits, migration, admixture, and bottlenecks specified in a template (.tpl) file — for large genomic regions and many independent loci. Its principal use, however, is inference: by comparing the observed (multi-dimensional) site frequency spectrum (SFS) of real data to spectra expected under a model, it uses a composite-likelihood approach with a conditional-expectation/maximization optimization to estimate demographic parameters and to compare competing models via likelihood/AIC. Inputs are a model template (.tpl), an estimation/parameter-range file (.est), and the observed SFS; outputs are maximum-likelihood parameter estimates and expected SFS across optimization runs. It is typically run with many independent starts to find the global optimum, and is a standard tool for demographic model testing in evolutionary genomics. | Population Genetics, Coalescent Theory, Simulation, Demographic Modeling | Population Genetics | Simulation Software | Documentation, Uses, and more |
| faststructure | mcc | 1.0 | fastStructure is a population-genetics tool for inferring genetic population structure and per-individual ancestry proportions from large SNP genotype datasets, serving the same goal as the classic STRUCTURE/ADMIXTURE programs but scaling to genome-wide data through a fast variational Bayesian inference algorithm. Given a chosen number of ancestral populations K, it estimates admixture proportions for each individual and allele frequencies for each population, offering both a simple and a logistic prior for handling weakly informative data. It reads standard PLINK (.bed/.fam/.bim) or STRUCTURE-format genotype input and writes per-K output files containing the admixture (Q) and allele-frequency (P) matrices. A typical workflow runs the structure inference across a range of K values, then uses the bundled chooseK utility to identify the K that best explains structure and the distruct utility to produce the familiar stacked-bar ancestry plots. It is distributed here (version 1.0) as one conda environment within a large multi-app genomics container on Rocky 8. | Population Genetics, Structural Biology, Genomics | Genetics, Biological Sciences | Inference Tool | Documentation, Uses, and more |
| fastx_toolkit | mcc, lcc | 1.0 | The FASTX-Toolkit is a classic collection of command-line utilities for preprocessing raw FASTA/FASTQ short-read data in next-generation sequencing pipelines. It bundles small single-purpose programs including fastq_quality_filter and fastq_quality_trimmer for quality-based filtering/trimming, fastx_clipper for adapter removal, fastx_trimmer for fixed-position trimming, fastx_collapser for collapsing identical reads, fastx_reverse_complement, and fastx_quality_stats/fastq_quality_boxplot for QC summaries and plots. Tools read FASTQ/FASTA on stdin and write to stdout so they chain easily in shell pipelines. It supports both Phred+33 and legacy Phred+64 quality encodings. This is version 0.0.14; it is largely superseded by newer trimmers but remains useful for lightweight per-base operations. | Bioinformatics, Short-Reads, Preprocessing | Biology, Biological Sciences | Pre-Processing Tool | Documentation, Uses, and more |
| fatotwobit | mcc | 1.0 | Documentation, Uses, and more | ||||
| feems | mcc, lcc | 1.0 | FEEMS (Fast Estimation of Effective Migration Surfaces) is a population-genetics method from the Novembre Lab for visualizing spatial population structure and heterogeneity in gene flow across a landscape. It models genetic data on a dense triangular grid laid over a geographic region and fits a Gaussian Markov random field whose edge weights represent effective migration rates, producing a color-coded migration surface where warm colors mark barriers to gene flow and cool colors mark corridors. It extends the classic EEMS idea but is dramatically faster because it solves a penalized-likelihood optimization rather than running a long MCMC, with a smoothness penalty (lambda) tuned by cross-validation. Inputs are typically a PLINK-format genotype dataset plus sample coordinates and a spatial grid/outer boundary; the Python API (a scikit-learn-style estimator) returns fitted edge weights that are then plotted with cartopy/matplotlib over a map. It is commonly used to complement PCA and admixture analyses when researchers want to see where geography structures genetic variation. | Documentation, Uses, and more | |||
| feh | mcc | N/A | feh is a lightweight and fast image viewer for the X Window System, designed for simplicity and speed. | Image Viewer, Lightweight, Fast, X11 | Software Engineering, Other Computer and Information Sciences | Application | Documentation, Uses, and more |
| fenics | lcc | N/A | FEniCS is a popular open-source computing platform for solving partial differential equations (PDEs) with the finite element method (FEM). FEniCS enables users to quickly translate scientific models into efficient finite element code. | partial differential equations, finite element method, numerical simulation, scientific computing, Python, C++ | Computational Mathematics, Computational Science | Scientific computing software | Documentation, Uses, and more |
| fenicsx | mcc | 1.0 | FEniCSx (whose computational core is DOLFINx) is an open-source platform for solving partial differential equations with the finite element method, letting users express variational (weak) formulations in a near-mathematical high-level syntax and have optimized C++ kernels generated and compiled automatically. Problems are written in Python (or C++) using the Unified Form Language (UFL) to declare function spaces, trial/test functions, and bilinear/linear forms; DOLFINx then handles mesh management, degree-of-freedom distribution, matrix/vector assembly, and boundary conditions, scaling across cores and nodes via MPI. It supports a wide range of element families and orders, mixed elements, and complex geometries, and integrates with PETSc (and SLEPc for eigenproblems, MUMPS for direct solves) for parallel linear and nonlinear solvers. This build is version 0.10.0.post4 compiled with Spack including py-fenics-dolfinx with petsc4py and slepc4py, on CentOS 8 with an OpenMPI+UCX/InfiniBand and PETSc/SLEPc/MUMPS toolchain tuned for distributed-memory HPC PDE simulations. Results are typically written to XDMF/VTX files for visualization in ParaView. | Finite Element Method, Partial Differential Equations, Computational Science, Open Source | Applied Mathematics, Mathematics | Library | Documentation, Uses, and more |
| fenicsx-mpc | mcc | 1.0 | dolfinx_mpc is an extension library for the FEniCSx/DOLFINx finite-element framework that adds support for multi-point constraints (MPCs) in variational (PDE) problems. Multi-point constraints express linear relationships between degrees of freedom, enabling boundary conditions and coupling that DOLFINx cannot express with standard Dirichlet conditions alone — most notably periodic boundary conditions, but also slip/contact-type conditions and general master-slave DOF relations. It works by assembling the finite-element system into the constrained solution space, providing custom assemblers for bilinear and linear forms defined with UFL, and integrates with DOLFINx's MPI-parallel PETSc-based linear algebra. Users write Python (or C++) scripts that build a MultiPointConstraint object from geometric or topological relations, then use dolfinx_mpc's constrained assembly routines in place of the standard DOLFINx ones. This build is from the dolfinx_mpc GitHub repository at tag v0.10.5, matched to a Spack-built FEniCSx 0.10.0 base. | Documentation, Uses, and more | |||
| ffmpeg | mcc | 1.0 | FFmpeg (conda-forge 8.0.0) is the de facto standard, cross-platform command-line framework for recording, converting, and streaming audio and video, built on the libav* libraries (libavcodec, libavformat, libavfilter, libswscale, and others) that decode and encode a vast range of media formats and codecs. It can transcode between formats, remux containers, resize/crop/pad and filter video, adjust audio, extract or overlay streams, generate thumbnails, concatenate and trim clips, capture from devices, and stream over network protocols, all through a powerful filtergraph system. In scientific/HPC contexts it is most often used to assemble image sequences into movies (e.g., rendering simulation or microscopy frames into an MP4), extract frames from video, or convert media for presentations. The core tools are ffmpeg (transcode/process), ffprobe (inspect media metadata and stream info), and ffplay (playback). Its capabilities are controlled by an extensive set of input/output options and video/audio filter chains. | Media Processing, Multimedia, Video, Audio | Computer Science | Tool | Documentation, Uses, and more |
| fftw | mcc, lcc | N/A | A Fast Fourier Transform library | Signal Processing, Data Compression, Partial Differential Equations | Applied Mathematics, Mathematics | Computational Software | Documentation, Uses, and more |
| filtlong | mcc | 1.0 | Filtlong is a quality-filtering tool for long sequencing reads (Oxford Nanopore and PacBio) that selects the best subset of reads to reduce data volume and improve downstream assembly, correction, or polishing. It scores reads by length and by quality (using per-base Phred quality when available, and optionally external Illumina short reads as a reference for k-mer-based quality assessment), then keeps reads according to user-set targets such as a total base budget, a percentage to keep, or minimum length/mean-quality thresholds. It reads a FASTQ (optionally gzipped) and streams filtered FASTQ to standard output. This is version 0.2.1. | Bioinformatics, Long Read Sequencing, Sequence Data Processing | Biological Sciences | Data Filtering Tool | Documentation, Uses, and more |
| finder | lcc | 1.0 | Finder software. | Documentation, Uses, and more | |||
| findutils | mcc, lcc, ecc | N/A | The GNU Find Utilities are the basic directory searching utilities of the GNU operating system. These programs are typically used in conjunction with other programs to provide modular and powerful directory search and file locating capabilities to other commands. | Unix, Linux, file searching, command-line tools, system utilities, shell scripting | Systems Programming, Computer Science | System utility software | Documentation, Uses, and more |
| finestructure | mcc | 1.0 | fineSTRUCTURE (version 2.1.3), together with its companion ChromoPainter, infers fine-scale population structure from dense genetic data such as phased SNP haplotypes. ChromoPainter first "paints" each individual's genome as a mosaic of chunks copied from all other sampled individuals, producing a coancestry (chunk-count) matrix that captures haplotype sharing; fineSTRUCTURE then applies a Bayesian MCMC model-based clustering to that matrix to group individuals into populations and build a tree of their relationships, resolving much finer structure than allele-frequency methods like STRUCTURE or PCA. Inputs are phased haplotype and recombination-map files; outputs are the coancestry matrix, cluster assignments, and a dendrogram, which are explored in the accompanying GUI or R tools. It is widely used in human and non-model population genomics to detect subtle ancestry patterns, migration, and admixture. A typical analysis chains ChromoPainter over chromosomes and then runs the fineSTRUCTURE MCMC and tree-building stages on the combined coancestry matrix. | population genetics, haplotype analysis, genetic clustering, ancestry inference, ChromoPainter, genomics | Statistical Genetics, Bioinformatics | Analysis software | Documentation, Uses, and more |
| firefox | mcc | N/A | Firefox is a free, open source Web browser for Windows, Linux and Mac OS X. | Web Browser, Internet, Privacy, Open-Source, Cross-Platform | Software Engineering, Computer & Information Sciences | Web Browser | Documentation, Uses, and more |
| fixchr | mcc | 1.0 | fixchr is a preprocessing utility from the Schneeberger lab that prepares two genome assemblies for structural comparison by identifying homologous chromosomes between them and reorienting them consistently. Comparative genomics and structural-variant tools (such as SyRI) require that the query and reference assemblies have matching chromosome identity and the same strand orientation; fixchr automates this by aligning the assemblies, selecting the correct set of homologous chromosomes, and flipping/reverse-complementing sequences as needed so the two genomes are in a comparable coordinate frame. It takes the two genome FASTA files (typically along with a whole-genome alignment) and outputs corrected, orientation-fixed assemblies ready for downstream synteny and rearrangement analysis. It is installed here directly from its GitHub repository. | Documentation, Uses, and more | |||
| fixesproto | mcc | N/A | Documentation, Uses, and more | ||||
| flash | mcc | 1.0 | FLASH (Fast Length Adjustment of SHort reads, version 1.2.11 from bioconda) merges overlapping paired-end reads into longer single reads by detecting and stitching the overlap between the forward and reverse mates. It is most useful when the DNA fragment (insert) is shorter than the combined read length, as with amplicon sequencing (e.g., 16S rRNA regions) or short-insert libraries, where merging recovers the full fragment sequence and improves downstream accuracy. Users control merging with minimum/maximum overlap and mismatch-ratio parameters and can supply an expected fragment length and read length. It produces merged (extended fragment) reads, the still-uncombined read pairs, and a histogram of merged lengths. It is a widely used preprocessing step ahead of amplicon clustering/OTU/ASV pipelines and short-insert assembly. | Computational Software, Numerical Simulation, Fluid Dynamics, Electromagnetics, Structural Analysis | Physical Sciences | Computational Physics | Documentation, Uses, and more |
| flex | mcc, lcc | N/A | flex (fast lexical analyzer generator) is a tool for generating scanners: programs which recognize lexical patterns in text. Description Source: https://github.com/westes/flex |
Lexical Analyzer, Text Processing, Pattern Matching, Scanner Generator | Software Engineering, Other Computer and Information Sciences | Compiler Tool | Documentation, Uses, and more |
| fltk | mcc, lcc | N/A | FLTK is a cross-platform C++ GUI toolkit for UNIX®/Linux® (X11), Microsoft® Windows®, and macOS®. FLTK provides modern GUI functionality without the bloat and supports 3D graphics via OpenGL® and its built-in GLUT emulation. Description Source: https://www.fltk.org/ |
Gui Toolkit, Cross-Platform Development, C++ Library | Computer Science, Engineering & Technology | Development Tool | Documentation, Uses, and more |
| fluent | lcc | N/A | Ansys Fluent (release 24R1) is an industry-leading commercial computational fluid dynamics (CFD) solver for modeling fluid flow, turbulence, heat and mass transfer, chemical reactions, combustion, and multiphase phenomena. It offers a broad range of physical models, meshing flexibility, and parallel scalability for engineering simulation across aerospace, automotive, energy, and process industries. Engineers use it to analyze and optimize designs involving fluid dynamics and thermal management. | CFD, fluid dynamics, heat transfer, engineering, simulation | Engineering, Mechanical Engineering | Commercial | Documentation, Uses, and more |
| flye | mcc, lcc | 1.0 | Flye is a de novo assembler for single-molecule long reads from PacBio and Oxford Nanopore platforms, capable of assembling genomes ranging from small bacteria to large plants and human, as well as metagenomes. It uses a repeat-graph approach: rather than requiring accurate reads, it builds assemblies from noisy long reads and resolves repeats using the structure of a repeat graph, then polishes the consensus. Input is one or more read files, with distinct modes for PacBio raw, PacBio HiFi, Nanopore raw, Nanopore corrected, and Nanopore high-quality data depending on chemistry and error profile, and output is a contigs/scaffolds FASTA plus an assembly graph (assembly_graph.gfa/.gv) and per-contig statistics. Flye also includes the metaFlye mode for uneven-coverage metagenomic data, accepts an optional genome-size hint, and reports contig circularity and multiplicity, making it a common first step in bacterial and long-read genome assembly pipelines. This is version 2.9.5. | Genome Assembly, Long-Read Sequencing, Nanopore Sequencing | Genomics, Biological Sciences | Bioinformatics Tool | Documentation, Uses, and more |
| font-util | mcc | N/A | X.Org font package creation/installation utilities and fonts. | font, typography, font management, font conversion, command-line tool, font subsetting | "", Software Engineering, Systems, and Development | Utility | Documentation, Uses, and more |
| fontconfig | mcc | N/A | Fontconfig is a library for configuring and customizing font access. Description Source: https://www.freedesktop.org/wiki/Software/fontconfig/ |
Fontconfig, Font Management, Font Configuration, Linux, Unix-Like Systems | Software Engineering, Engineering & Technology | System Tool | Documentation, Uses, and more |
| fontsproto | mcc | N/A | Documentation, Uses, and more | ||||
| fragpipe | mcc | 1.0 | FragPipe is an integrated computational proteomics pipeline, offered as a Java graphical application with an accompanying headless/command-line mode, that ties together a suite of tools for analyzing mass-spectrometry-based proteomics and peptidomics data. At its core is MSFragger, an ultrafast database-search engine, orchestrated together with Philosopher (for target-decoy validation, PeptideProphet/ProteinProphet-style FDR filtering, and protein inference), IonQuant (label-free and isobaric quantification with match-between-runs), TMT-Integrator (for TMT/iTRAQ isobaric-labeling quantitation), and Crystal-C/PTM-Shepherd for open/PTM searches. It supports closed, open, and offset mass searches enabling discovery of post-translational modifications, as well as glyco and DIA workflows. Users typically configure a workflow (choosing a built-in template such as LFQ, TMT, or Open Search), provide raw spectra (e.g. mzML, or vendor formats via the bundled converters) and a protein FASTA, and FragPipe runs identification, validation, and quantification end-to-end, producing PSM/peptide/protein tables. This deployment packages FragPipe 20.0, derived from the upstream fcyucn/fragpipe Docker image and relocated to /opt. | Proteomics, Mass Spectrometry, Data Analysis, Bioinformatics | Biochemistry and Molecular Biology, Proteomics | Data Analysis Tool | Documentation, Uses, and more |
| freebayes | mcc, lcc | 1.0 | FreeBayes is a Bayesian, haplotype-based genetic variant detector that identifies SNPs, indels, multi-nucleotide polymorphisms, and complex composite events from aligned short-read sequencing data. Rather than evaluating each position independently, it uses the literal read sequences spanning a region to infer the most likely set of haplotypes and their allele frequencies, making it robust for calling clustered variants and for arbitrary ploidy and pooled samples. It reads one or more coordinate-sorted, indexed BAM files together with the reference FASTA and emits variants in standard VCF, with options to set ploidy, pooled/continuous allele-frequency modes, minimum mapping/base qualities, and region targeting. It is frequently used for population and non-diploid variant calling. This build is the prebuilt static Linux binary freebayes 1.3.4. | Genetic Variant Detector, Snp Detection, Indel Detection, Bayesian Algorithm | Bioinformatics, Biological Sciences | Tool | Documentation, Uses, and more |
| freeglut | mcc | N/A | freeglut is a free-software/open-source alternative to the OpenGL Utility Toolkit (GLUT) library. GLUT (and hence freeglut) takes care of all the system-specific chores required for creating windows, initializing OpenGL contexts, and handling input events, to allow for trully portable OpenGL programs. Description Source: https://freeglut.sourceforge.net/ |
Graphics Programming, Opengl, Cross-Platform Development, Free Software | Computer Science, Computer & Information Sciences | Graphics Library | Documentation, Uses, and more |
| freetype | mcc, lcc | N/A | FreeType is a freely available software library to render fonts. Description Source: https://freetype.org/ |
Font Engine, Text Rendering, Typography | Software Engineering, Computer & Information Sciences | Font Engine | Documentation, Uses, and more |
| fribidi | mcc | N/A | The Free Implementation of the Unicode Bidirectional Algorithm. Description Source: https://github.com/fribidi/fribidi?tab=readme-ov-file |
Unicode, Bidirectional Algorithm | Natural Language Processing, Computer & Information Sciences | Text Processing | Documentation, Uses, and more |
| funannotate | mcc | 1.0 | funannotate is a comprehensive genome annotation pipeline optimized for fungal (and small eukaryotic) genomes, covering the path from assembly to submission-ready annotation. It performs assembly cleanup and sorting, repeat soft-masking, RNA-seq-informed and ab initio gene prediction (training and combining Augustus, GeneMark, snap, and GlimmerHMM via EVidenceModeler), tRNA prediction, and functional annotation by integrating InterProScan, Pfam, dbCAN (CAZymes), MEROPS, SwissProt, and antiSMASH results. It also includes a comparative-genomics module and produces GenBank/GFF3 outputs plus annotation reports. The workflow proceeds through subcommands that clean, sort, and mask the assembly, train and predict gene models, update annotations with RNA-seq evidence, and add functional annotation. This is version 1.8.17. | Genome Prediction, Genome Annotation, Fungi, Bioinformatics | Biological Sciences | Bioinformatics Tool | Documentation, Uses, and more |
| fuse-overlayfs | ecc | N/A | This project provides a convenient way to automatically perform a static build using a container. The result is a self-contained binary without dependencies, that can be copied across hosts. | FUSE,Overlay Filesystem,Userspace Filesystem,Linux | Systems Engineering, Computer Science | System utility software | Documentation, Uses, and more |
| g_kuh | mcc | 1.0 | Trajectory analysis tool contributed by the SMOG structure-based model extension to GROMACS 4.5.4, which computes native contact fraction (Q) along a trajectory. Q is the standard reaction coordinate for structure-based folding studies, measuring how many of a reference structure contacts are formed in each frame, so that folding and unfolding transitions can be located and free energy profiles built. It reads a trajectory together with the run input describing the contact set and writes a time series suitable for further histogramming. This tool exists only in SMOG-patched GROMACS builds and is absent from standard distributions. | Documentation, Uses, and more | |||
| galba | mcc | 1.0 | GALBA (packaged from the katharinahoff/galba-notebook Docker image) is a pipeline for structural genome annotation of eukaryotes that trains and runs the AUGUSTUS gene predictor using evidence from homologous proteins, making it a member of the BRAKER family of annotation tools. It is specifically designed for the common situation where RNA-seq is unavailable but protein sequences from one or a few closely related species are: it aligns those proteins to the target genome (using miniprot or GenomeThreader), derives training gene structures, trains AUGUSTUS, and produces genome-wide gene predictions. It is particularly well suited to large genomes and to cases where using many species' proteins would degrade accuracy, since GALBA emphasizes closely related protein evidence. Inputs are a soft-masked genome assembly and a protein FASTA; the output is a set of predicted gene models in GTF/GFF3. The notebook-based image bundles AUGUSTUS and the required aligners and dependencies for a ready-to-run annotation environment. | Documentation, Uses, and more | |||
| gamess | mcc, lcc | 1.0 | General Atomic and Molecular Electronic Structure System | Quantum Chemistry, Electronic Structure Calculations, Ab Initio Calculations | Physical Sciences, Chemical Sciences | Quantum Chemistry | Documentation, Uses, and more |
| gangstr | mcc | 1.0 | GangSTR is a specialized genotyper for short tandem repeats (STRs) that, unlike most STR callers, is designed to genotype repeats genome-wide including long expansions that exceed the sequencing read length. It jointly analyzes multiple sources of evidence from short-read whole-genome sequencing (enclosing reads, spanning read pairs, fully-repetitive reads, and flanking reads) within a maximum-likelihood framework to estimate the diploid repeat copy number at each locus. Inputs are an indexed BAM/CRAM alignment, a reference FASTA, and a BED-like reference panel of repeat regions with their motifs; output is a VCF annotating each locus with the two allele lengths and supporting-evidence details. It is frequently used in studies of repeat-expansion disorders and is often paired with tools like dumpSTR for filtering the resulting calls. | Str, Genotyping, Short Tandem Repeats, Sequencing Data Analysis | Genomics, Biological Sciences | Genomic Analysis Tool | Documentation, Uses, and more |
| gapless-0.4 | mcc | 1.0 | gapless is a genome-assembly-improvement tool that closes gaps and increases contiguity in draft assemblies using long reads (PacBio or Oxford Nanopore). It works iteratively: mapping long reads to the current assembly, identifying scaffold gaps and breakpoints, and using spanning reads to fill Ns, extend contig ends, and resolve or correct problematic joins, converging over successive rounds toward a more complete assembly. Inputs are an assembly FASTA and long-read data (plus alignments produced with a mapper such as minimap2), and outputs are the gap-filled, extended assembly along with logs of the modifications made. It is typically run as part of a long-read assembly-finishing pipeline after initial contig assembly and scaffolding. This is bioconda gapless version 0.4. | Documentation, Uses, and more | |||
| gatk | mcc, lcc | 1.0 | GATK tools . | Variant Calling, Genotyping, Sequencing Analysis | Biology, Biological Sciences | Tool | Documentation, Uses, and more |
| gatk4 | mcc | 1.0 | GATK4 (Genome Analysis Toolkit), version 4.2.2.0, from the Broad Institute, is the leading toolkit for variant discovery in high-throughput sequencing data, covering germline SNP/indel calling, somatic (tumor) SNV/indel and copy-number calling, and the surrounding data-preparation steps. It follows the GATK Best Practices workflow: mark duplicates and base-quality score recalibration (BaseRecalibrator/ApplyBQSR), per-sample calling with HaplotypeCaller in GVCF mode, joint genotyping across cohorts (GenomicsDBImport/GenotypeGVCFs), and variant filtering (VQSR or hard filters); somatic analysis uses Mutect2. GATK4 is distributed as a single gatk launcher wrapping many tools, and several tools are implemented on Apache Spark for parallel scaling. It reads BAM/CRAM alignments plus a reference and known-sites resources and produces VCF/GVCF output, and is the de-facto standard for human and non-human germline/somatic variant analysis. | Bioinformatics, Genomic Analysis, Variant Discovery, High-Throughput Sequencing | Biological Sciences | Genomic Analysis Tool | Documentation, Uses, and more |
| gaussian | mcc, lcc | N/A | Gaussian is a computational chemistry software for quantum mechanical simulations, widely used by researchers for studying molecular structures and reactions. It offers advanced capabilities in electronic structure prediction and various spectroscopic properties analysis. | Computational Chemistry, Quantum Chemistry, Molecular Modeling | Chemical Sciences, Natural Sciences | Simulation Software | Documentation, Uses, and more |
| gawk | mcc, ecc | N/A | Gawk is the GNU project's version of the AWK programming language. It is a standard Linux tool for text processing, data extraction, and quick reporting. | Text processing, Scripting, Data analysis | "", Computer Science | Command-line tool | Documentation, Uses, and more |
| gbs-snp-crop | lcc | 1.0 | GBS-SNP-CROP (GBS SNP Calling Reference Optional Pipeline) is a Perl-based pipeline for calling SNPs and performing reference-free (or reference-based) genotyping from genotyping-by-sequencing (GBS) and low-coverage sequencing data, designed for population and breeding studies in non-model species. It runs as a sequence of numbered stages that parse and demultiplex reads, trim with Trimmomatic, merge overlapping paired reads with PEAR, build a mock reference by clustering with VSEARCH, align reads with BWA and process alignments with SAMtools, and finally call and filter variants into a genotype matrix. It is executed as successive GBS-SNP-CROP Perl scripts sharing common parameters, and this build (v4.1) is packaged with a conda environment providing the required Perl modules and aligner dependencies. | genotyping-by-sequencing, SNP filtering, variant curation, linkage mapping, population genetics, plant genomics | Genetic Mapping, Bioinformatics | Analysis software | Documentation, Uses, and more |
| gcc | mcc | N/A | GCC, the GNU Compiler Collection, is an open-source compiler system that supports various programming languages like C, C++, Objective-C, Fortran, Ada, and Go. It is widely used in software development for compiling source code into executable programs, providing robust performance and extensive compatibility across different platforms and operating systems. \r Description Source: https://gcc.gnu.org/ |
Compiler, Software Development, Programming | Computer Science | Development Tools | Documentation, Uses, and more |
| gcc-runtime | ecc | N/A | The GCC Runtime Library is a part of the GCC compiler collection that provides runtime support for programs compiled with GCC. It includes various runtime libraries and support files necessary for programs compiled with GCC to execute correctly. | Compiler, Runtime, Library | Software Engineering, Systems & Development, Engineering & Technology | Runtime Library | Documentation, Uses, and more |
| gcta | lcc | 1.0 | GCTA (Genome-wide Complex Trait Analysis) is a tool for estimating the genetic architecture of complex traits from genome-wide SNP data, best known for the GREML method that partitions phenotypic variance into components explained by all SNPs to estimate SNP-based heritability. Beyond variance-component estimation, it computes genomic-relationship matrices (GRMs), estimates genetic correlations between traits (bivariate GREML), performs mixed-linear-model association analysis (MLMA) and conditional/joint GWAS analysis (COJO), runs principal-component analysis, and offers simulation and LD-based tools. It operates on PLINK binary genotype files (.bed/.bim/.fam); a typical two-step workflow first builds a genomic-relationship matrix and then fits the REML model with phenotype and covariate files. This is the version 1.95.1 Linux binary. GCTA is a foundational package in statistical and quantitative genetics for dissecting the heritability and genetic basis of complex traits and diseases. | genetics, bioinformatics, GWAS, heritability, SNP analysis | Bioinformatics, Genetics | Analysis Tool | Documentation, Uses, and more |
| gdal | mcc, lcc | 1.0 | GDAL (Geospatial Data Abstraction Library), version 3.10.3, is the foundational open-source translator library and command-line toolset for raster and vector geospatial data, underpinning most GIS and Earth-observation software. It provides a single abstract data model with drivers for hundreds of formats — GeoTIFF, NetCDF, HDF, JPEG2000, COG, Shapefile, GeoPackage, GeoJSON, and many more — and handles coordinate reference systems, reprojection, resampling, and format conversion. Its widely used utilities include gdal_translate (format/subset conversion), gdalwarp (reprojection, mosaicking, resampling), gdalinfo/ogrinfo (metadata inspection), gdal_rasterize/gdal_polygonize (raster-vector conversion), and ogr2ogr (vector translation and filtering). It exposes C/C++ and Python (osgeo) APIs used by tools like QGIS, rasterio, and PostGIS. In a research setting it is the workhorse for preprocessing satellite imagery, DEMs, and spatial layers into analysis-ready form. | Geospatial Data, Raster Data, Vector Data, Data Formats, Geospatial Analysis | Geographic Information Systems, Earth & Environmental Sciences, Engineering & Technology | Library | Documentation, Uses, and more |
| gdb | mcc | N/A | GDB, the GNU Project debugger, allows you to see what is going on `inside' another program while it executes -- or what another program was doing at the moment it crashed. Description Source: https://sourceware.org/gdb |
Debugger, Programming, Software Development | Computer Science, Computer & Information Sciences | Debugger | Documentation, Uses, and more |
| gdbm | mcc, lcc, ecc | N/A | A library of database functions that use extensible hashing and works similar to the standard UNIX dbm functions. Description Source: https://github.com/conda-forge/gdbm-feedstock/blob/main/recipe/meta.yaml | Database Management, Key-Value Store, Persistent Storage, Api | Computer Science, Computer & Information Sciences | Library | Documentation, Uses, and more |
| gdk-pixbuf | mcc | N/A | The Gdk Pixbuf package is a toolkit for image loading and pixel buffer manipulation. It is used by GTK+ 2 and GTK+ 3 to load and manipulate images. In the past it was distributed as part of GTK+ 2 but it was split off into a separate package in preparation for the change to GTK+ 3. Description Source: https://www.linuxfromscratch.org/blfs/view/11.2/x/gdk-pixbuf.html |
Image Processing, Graphics, Library | Computer Science, Computer & Information Sciences | Image Processing | Documentation, Uses, and more |
| geant4 | mcc, lcc | 1.0 | Geant4 is a large C++ toolkit for Monte Carlo simulation of the passage of particles through matter, developed at CERN and used extensively in high-energy, nuclear, accelerator, space, and medical physics (including radiotherapy and detector design). It models the transport and interactions of particles across a wide energy range through user-defined 3D detector geometries and materials, offering comprehensive physics processes (electromagnetic, hadronic, decay, optical) selectable via physics lists, plus tracking, event/run management, scoring, and visualization. Users build applications against the Geant4 libraries, defining geometry, primary particle sources, and sensitive detectors, and control runs through macro command scripts. This is version 11.3.2 from conda-forge; the surrounding container also bundles ROOT for analysis of the simulated hit/energy-deposition data. | Software, Physics, Simulation, Particle Transport | Particle Physics, Physical Sciences | Toolkit | Documentation, Uses, and more |
| gemelli | mcc | 1.0 | Gemelli is a QIIME 2 plugin for compositional dimensionality reduction of microbiome data, from the Knight and Shenhav labs. Its distinguishing capability is Compositional Tensor Factorization, which handles repeated measures designs where the same subject is sampled at multiple timepoints or body sites, separating variation between subjects from variation across time or site in a way that ordinary principal coordinates analysis cannot. For cross sectional studies with a single sample per subject it provides Robust Aitchison PCA, a sparse compositional ordination tolerant of the many zeros typical of count tables, and TEMPTED handles longitudinal designs with irregular sampling times. Phylogeny-aware variants report loadings over tree branches as well as features, joint factorization integrates several tables at once for multi-omics work, and a rarefaction quality-control check reports whether rarefying is distorting the resulting ordination. Inputs are a feature table plus sample metadata naming the subject identifier and the state column, and outputs are ordinations, sample and feature loadings, and distance matrices that feed directly into the rest of the QIIME 2 ecosystem. | Documentation, Uses, and more | |||
| genmap | mcc | 1.0 | GenMap is a fast, memory-efficient tool for computing genome mappability and uniqueness over large genomes, part of the SeqAn-based ecosystem. Mappability quantifies, for every position in a genome, how uniquely a k-mer starting there maps back to the genome allowing up to a set number of mismatches, which is essential for masking repetitive regions, weighting read-depth signals, and designing probes or CRISPR guides. Its workflow has two stages: first an indexing step builds an FM-index from a reference FASTA, then a mapping step computes per-position mappability for a chosen k-mer length and mismatch allowance and exports it in formats such as raw, WIG, BEDGRAPH, or plain-text (tab-separated) tracks. This is bioconda genmap 1.3.0. | Genetic Mapping, Linkage Analysis, Genetic Loci Mapping, Family-Based Studies | Biological Sciences | Genetic Mapping Software | Documentation, Uses, and more |
| genomad | mcc | 1.0 | geNomad (version 1.11.1 from bioconda) identifies mobile genetic elements, specifically virus (including proviruses/prophages) and plasmid sequences, in genomic and metagenomic assemblies, and assigns taxonomy to the detected viruses. It combines a gene-based classifier that uses a large marker/annotation dataset of virus- and plasmid-specific proteins with an alignment-free neural-network classifier, and merges the two branches for robust scoring; it can also excise integrated proviruses from host contigs. Its end-to-end mode runs annotation, marker classification, the neural-network branch, aggregated classification, and (for viruses) taxonomic assignment, producing per-contig virus and plasmid summaries with scores and predicted taxonomy. It requires a downloadable geNomad database. It is a leading tool for viral and plasmid discovery in metagenomes and has largely supplanted older single-purpose virus finders in many pipelines. | Documentation, Uses, and more | |||
| genomescope2 | mcc | 1.0 | GenomeScope2 is a reference-free tool for estimating key genome characteristics directly from sequencing reads, widely used to profile a genome before assembly. It fits a mixture model of negative binomial distributions to a k-mer frequency histogram (typically generated by Jellyfish or KMC), from which it infers genome size, overall heterozygosity, repeat content, and read error rate; version 2 adds an explicit polyploid-aware model that can fit diploid, triploid, tetraploid, and higher ploidies, and it is commonly paired with the companion tool Smudgeplot (from the same authors) for independent ploidy estimation. Input is a k-mer histogram file plus the k-mer length and (for polyploids) the ploidy, and output is a set of fitted-model plots (linear and log) along with a text summary of parameter estimates and confidence intervals, making it a standard first QC step for evaluating heterozygosity and sizing assemblies. This is bioconda genomescope2 2.0.1. | Genome Analysis, Bioinformatics | Genomics, Biological Sciences | Bioinformatics Tool | Documentation, Uses, and more |
| genometools | mcc | 1.0 | GenomeTools is a free collection of bioinformatics utilities and a C library, unified under the single 'gt' binary, for processing and analyzing genome annotations and sequences; note that in this container it is provided via the Dfam TE Tools v1.9.5 image (dfam/tetools:1.95), where it accompanies the transposable-element toolchain. The gt binary exposes many subtools for working primarily with GFF3 annotation files and sequences, including a gff3 subtool that parses, validates, sorts, and tidies GFF3, format interconversion such as gff3_to_gtf, a sketch subtool that renders annotations to publication-quality PNG/PDF/SVG diagrams, extractfeat for extracting feature sequences, seqstat for sequence statistics, and specialized indexing/analysis tools such as suffix-array construction and the LTRharvest/LTRdigest tools for de novo LTR-retrotransposon discovery. It is valued for robust, standards-compliant GFF3 handling and is often used to validate and normalize annotation files before or after other pipeline steps. | genomics, bioinformatics, sequence analysis, data visualization | Bioinformatics, Genomics | Open Source | Documentation, Uses, and more |
| getorganelle | lcc | 1.0 | GetOrganelle is a toolkit for de novo assembly of organelle genomes—chloroplast (plastome), mitochondrial, and nuclear ribosomal DNA—directly from whole-genome sequencing reads. It works by recruiting organelle-derived reads through iterative baiting against seed reference sequences, then assembling them with SPAdes and using its own graph-based algorithm to untangle the assembly graph and export the correct circular (or linear) organelle genome, including handling of the chloroplast inverted-repeat structure. The main workflow specifies read files, an organelle type, and an output directory, producing one or more complete organelle genome FASTA files plus the assembly graph for inspection in Bandage. Version 1.7.5.0 is a de facto standard in plant plastome and organellar phylogenomics projects because of its high success rate at recovering complete, correctly structured organelle genomes. | Bioinformatics, Genomics, Molecular Biology, Computational Biology | Genomics, Biological Sciences | Analysis Tool | Documentation, Uses, and more |
| gettext | mcc, lcc, ecc | N/A | Gettext is a software framework used for internationalization and localization of software applications, allowing them to be translated into multiple languages. It provides tools for developers to extract strings from their code, create and manage translation files, and compile these files into formats that can be used by the application to display the translated text, thereby making the software accessible to a global audience. | Internationalization, Localization, Software Development | Software Engineering, Computer & Information Sciences | Localization Tool | Documentation, Uses, and more |
| gfacpp | mcc | 1.0 | Documentation, Uses, and more | ||||
| gfclient | mcc | 1.0 | Documentation, Uses, and more | ||||
| gff3toolkit | lcc | 1.0 | GFF3toolkit (installed as the gff3tool PyPI package, from the USDA NAL-i5K project) is a Python toolkit for quality-controlling and manipulating GFF3 genome-annotation files. Its flagship utility, gff3_QC, checks a GFF3 against roughly fifty documented error types — such as inconsistent parent/child feature coordinates, wrong phase on CDS features, features extending beyond their parent, and duplicate IDs — reporting problems so annotations can be fixed before submission or downstream use. Companion commands include gff3_fix (automatically corrects many of the errors flagged by gff3_QC), gff3_sort (topologically sorts features so children follow parents), gff3_merge (combines/replaces annotations from two GFF3 files), gff3_to_fasta (extracts spliced/CDS/protein/trans-spliced sequences given a genome FASTA), and gff3_ID_generator. A typical workflow runs the QC check against an annotation and genome and then applies the fix utility to repair the reported issues. It is widely used for cleaning up eukaryotic genome annotations, particularly in arthropod/i5K projects. | Documentation, Uses, and more | |||
| gfortran | mcc | N/A | GNU Fortran is the Fortran compiler of the GNU Compiler Collection (GCC). It supports Fortran 77, Fortran 90, Fortran 95, Fortran 2003, Fortran 2008, and Fortran 2018 standards. | Fortran, Compiler, Open Source, GCC | Numerical Methods, Applied Mathematics | Compiler | Documentation, Uses, and more |
| gfserver | mcc | 1.0 | Documentation, Uses, and more | ||||
| ghostscript | mcc | N/A | Ghostscript is an interpreter for the PostScript® language and PDF files. It is available under either the GNU GPL Affero license or licensed for commercial use from Artifex Software, Inc. Ghostscript consists of a PostScript interpreter layer and a graphics library. Description Source: https://ghostscript.com/ |
Document-Processing, File-Conversion, Print-Management | Computer Science, Computer & Information Sciences, Engineering & Technology | Software Development | Documentation, Uses, and more |
| ghostscript-fonts | mcc | N/A | Ghostscript fonts are a collection of fonts used by Ghostscript, a suite of software that provides an interpreter for the PostScript language and for PDF. | fonts, Ghostscript, PostScript, PDF | Computer Science, Computer and Information Sciences | Font Library | Documentation, Uses, and more |
| giflib | mcc | N/A | Giflib is a software package designed for reading and writing GIF (Graphics Interchange Format) images. It provides a set of utilities for processing GIFs, including conversion to and from various formats, making it a valuable tool for developers working with GIF animations and image processing in applications or web services. | Library, Image Processing, Gif Images | Image Processing, Computer Science | Package | Documentation, Uses, and more |
| gist-pipeline | lcc | 1.0 | The GIST pipeline (Galaxy IFU Spectroscopy Tool) is a Python framework for the systematic analysis of integral-field-unit (IFU) spectroscopic data of galaxies, where each spatial pixel of the field of view carries a full spectrum. It provides a modular, configurable pipeline that reads reduced IFU data cubes, spatially bins spaxels (typically via Voronoi tessellation to reach a target signal-to-noise), and then performs full-spectrum fitting to derive stellar and gas kinematics (velocity, velocity dispersion), stellar population properties (age, metallicity), and emission-line fluxes, largely built on the pPXF and GandALF fitting engines. It produces per-bin maps and tables of these physical quantities plus diagnostic plots, and its behavior is controlled through configuration files specifying modules and parameters. It is used in extragalactic astrophysics to characterize galaxy dynamics and stellar populations from instruments such as MUSE. This is the pip-installed gistPipeline package. | Documentation, Uses, and more | |||
| git | mcc, lcc | N/A | Git distributed version control (bundles OpenSSL 3.6 and libcurl 8.21) | Version Control, Software Development, Source Code Management | Computer Science, Computer & Information Sciences, Software Engineering, Systems & Development | Development Tools | Documentation, Uses, and more |
| git-lfs | lcc | N/A | git-lfs software. | Version Control, File Storage, Git Extension | Version Control, Engineering & Technology | Development Tool | Documentation, Uses, and more |
| gizmo | lcc | N/A | GIZMO is a widely used, massively parallel, multi-method code for astrophysical and cosmological simulations. Built here with the Intel compiler toolchain and MVAPICH2 MPI stack, it solves hydrodynamics via meshless finite-volume/mass methods coupled with self-gravity, and supports magnetohydrodynamics, radiative cooling, star formation, and feedback physics. It is used by researchers to model galaxy formation, star and planet formation, accretion disks, and large-scale structure. | Documentation, Uses, and more | |||
| gklib | mcc | 1.0 | A library of various helper routines and frameworks used by Karypis Lab software such as METIS. | Geometric Kernels, Computational Geometry, Algorithms | Applied Mathematics, Mathematics | Computational Software | Documentation, Uses, and more |
| gl2ps | mcc | N/A | GL2PS is a C library providing high quality vector output for any OpenGL application. The main difference between GL2PS and other similar libraries is the use of sorting algorithms capable of handling intersecting and stretched polygons, as well as non manifold objects. GL2PS provides advanced smooth shading and text rendering, culling of invisible primitives, mixed vector/bitmap output, and much more... Description Source: https://www.geuz.org/gl2ps/ |
Opengl, Postscript, Printing, Graphics, Vector Graphics | Visualization, Computer Science | Library | Documentation, Uses, and more |
| glib | mcc, ecc | N/A | GLib is a general-purpose, portable utility library, which provides many useful data types, macros, type conversions, string utilities, file utilities, a mainloop abstraction, and so on. It is one of the base libraries of GTK+. Description Source: https://docs.gtk.org/glib/ |
Utility Library, C Programming, Data Structures, Memory Management, Thread Support | Software Engineering, Computer & Information Sciences | Utility | Documentation, Uses, and more |
| glibc | ecc | N/A | glibc, the GNU C Library, is an essential part of most systems running the Linux kernel. It provides the necessary functionality for programs written in the C programming language to interact with the operating system and hardware. | C Library, Operating System, System Programming | Computer & Information Sciences | Library | Documentation, Uses, and more |
| glimpse | mcc | 1.0 | GLIMPSE2 is a method for genotype imputation and haplotype phasing specifically optimized for low-coverage whole-genome sequencing, where each site is covered by only a fraction of a read on average. Rather than relying on genotype calls, it works from genotype likelihoods (e.g. produced by bcftools mpileup) and a reference haplotype panel to statistically reconstruct accurate diploid genotypes across the genome, making low-coverage sequencing a cost-effective alternative to genotyping arrays. The workflow proceeds through discrete steps: splitting the reference panel into chunks, building binary reference files, imputing and phasing each chunk, and ligating the chunks back into whole chromosomes. Version 2.0.0 greatly improves speed and memory over GLIMPSE1 and adds support for very large reference panels, and it is commonly used in ancient-DNA and large-cohort low-pass sequencing studies. | Documentation, Uses, and more | |||
| glnexus | lcc | 1.0 | GLnexus (version 1.4.1) is a scalable system for joint variant calling that merges per-sample gVCFs into a unified, cohort-level multi-sample VCF/BCF. It performs the computationally demanding "joint genotyping" step — reconciling variant records across many samples, resolving overlapping and multi-allelic sites, and producing consistent genotypes and quality metrics — using an efficient database backend that lets it scale to tens or hundreds of thousands of samples on modest hardware. It ships with configuration presets tuned for the outputs of common callers, notably DeepVariant (as in the DeepVariant + GLnexus population workflow), as well as GATK-style gVCFs. Input is a set of gVCF files (and optionally a BED of target regions); output is a joint-genotyped BCF, typically piped to bcftools. It is a standard component of large-cohort resequencing and population-genomics pipelines. | Variant Calling, Genotyping, Haplotype-Aware, Bioinformatics | Biological Sciences | Bioinformatics Tool | Documentation, Uses, and more |
| glnexus-1.4.1 | mcc | 1.0 | GLnexus (GL, for genotype likelihood, nexus) is a scalable joint variant caller and gVCF merging engine designed for cohort-scale genomics. It takes per-sample gVCFs — typically produced by DeepVariant, GATK HaplotypeCaller, or similar callers — and performs joint genotyping across hundreds to hundreds of thousands of samples, resolving overlapping records and producing a unified project-level multi-sample VCF/BCF. Internally it uses an embedded RocksDB key-value store to manage the merge without holding everything in memory, enabling it to scale to very large cohorts on a single node. It is driven through configuration presets (e.g. DeepVariant, DeepVariantWGS, gatk) that tune the QC filtering and record-unification behavior for the upstream caller. Output is a bgzipped BCF that is normally piped through bcftools for downstream analysis. This build is version 1.4.1 from Bioconda. | genomics, bioinformatics, variant calling, high-performance computing | Bioinformatics, Genomics | Command-line tool | Documentation, Uses, and more |
| globus-cli | mcc, lcc | N/A | Globus CLI is the command-line interface to the Globus research data-management platform, provided on LCC as a Miniconda environment (module ccs/conda/globus-cli/3.29.0). It gives users the globus command for reliable, high-performance, fire-and-forget transfer and sharing of large datasets between Globus endpoints/collections, along with endpoint search, task management, and permission/ACL control. It is a widely used, user-facing HPC data-transfer utility that researchers invoke directly to move data in and out of the cluster. | Command-Line Interface, File Transfer Service, Data Management, Data Transfer | Computer Science, Other Computer & Information Sciences | Command Line Tool | Documentation, Uses, and more |
| glpk | mcc | N/A | The GLPK (GNU Linear Programming Kit) package is intended for solving large-scale linear programming (LP), mixed integer programming (MIP), and other related problems. It is a set of routines written in ANSI C and organized in the form of a callable library. Description Source: https://www.gnu.org/software/glpk/ |
Linear Programming, Optimization, Open-Source | Applied Mathematics, Mathematics | Mathematical Optimization | Documentation, Uses, and more |
| glproto | mcc, lcc | N/A | Documentation, Uses, and more | ||||
| gmake | mcc, ecc | N/A | GNU Make is a tool that controls the generation of executables and other non-source files from a program's source files. It is widely used for managing build processes in software development. | Build Automation, Makefile, Open Source, GNU | Computer Science, Software Engineering | Build System | Documentation, Uses, and more |
| gmap | lcc | 1.0 | GMAP/GSNAP (the gmap package, version 2021.08.25 from bioconda) is a genomic mapping and alignment suite from Thomas Wu. GMAP aligns longer cDNA, mRNA, and EST sequences to a reference genome with sensitivity to introns, making it well suited to spliced alignment and gene-structure work, while its companion GSNAP aligns short reads (including SNP- and methylation-aware modes and paired-end RNA-seq with splice detection). Users first build a genome index, then align transcript sequences with GMAP or run GSNAP for short-read spliced alignment. It outputs alignments in formats including SAM/BAM, GFF3, and PSL, and can report multiple alignment paths. It is widely used for transcript-to-genome alignment, splice-site discovery, and cross-species cDNA mapping. | Bioinformatics, Computational Biology, Genomics, Transcriptomics, RNA-Seq | Genomics, Biological Sciences | Bioinformatics Tool | Documentation, Uses, and more |
| gmp | mcc, lcc, ecc | N/A | GNU MP is a portable library written in C for arbitrary precision arithmetic on integers, rational numbers, and floating-point numbers. It aims to provide the fastest possible arithmetic for all applications that need higher precision than is directly supported by the basic C types. Description Source: https://gmplib.org/manual/Introduction-to-GMP |
Computational Software, Library, Mathematics, Software Development | Pure Mathematics, Mathematics | Library | Documentation, Uses, and more |
| gmsh | lcc | N/A | Mesh software . | mesh generation, finite element analysis, CAD, post-processing | Engineering, Finite Element Analysis | Mesh Generator | Documentation, Uses, and more |
| gnu | lcc | N/A | GNU Compiler Family (C/C++/Fortran for x86_64) | Open Source, Development Tools, Utilities | Software Engineering, Computer Science | Development Tools | Documentation, Uses, and more |
| gnu-parallel | mcc | 1.0 | GNU Parallel is a general-purpose command-line tool for executing shell jobs concurrently, letting a user saturate all available CPU cores (or distribute work across multiple machines over SSH) instead of running commands one at a time. It reads a list of inputs — from arguments, a file, or standard input — and runs a specified command once per input, controlling the degree of parallelism with a jobs option, substituting inputs into the command template with replacement strings (for the full input, its basename, dirname, or job number, and related forms), and preserving ordered output when requested. It supports building input combinations from multiple lists, retries, job resumption via a job log, progress reporting, and remote execution across SSH logins. In HPC and bioinformatics workflows it is a lightweight way to parallelize embarrassingly parallel per-file or per-sample tasks. This build is version 20240922 from conda-forge. | parallel computing, job scheduling, command-line tools, GNU, open-source | Computer Science | Command-line tool | Documentation, Uses, and more |
| gnu12 | lcc | N/A | GNU Compiler Family (C/C++/Fortran for x86_64) | Documentation, Uses, and more | |||
| gnu12/12.1.0 | mcc | N/A | GNU Compiler Family (C/C++/Fortran for x86_64) | Documentation, Uses, and more | |||
| gnu12/openmpi4 | mcc | N/A | A powerful implementation of MPI/SHMEM | Documentation, Uses, and more | |||
| gnu14 | lcc, ecc | N/A | GNU Compiler Family (C/C++/Fortran for x86_64) | Documentation, Uses, and more | |||
| gnu7 | lcc | N/A | GNU Compiler Family (C/C++/Fortran for x86_64) | Documentation, Uses, and more | |||
| gnu8 | lcc | N/A | GNU Compiler Family (C/C++/Fortran for x86_64) | Documentation, Uses, and more | |||
| gnu9 | mcc | N/A | GNU Compiler Family (C/C++/Fortran for x86_64) | Documentation, Uses, and more | |||
| gnu_parallel | mcc | N/A | GNU parallel 20240922 | parallel computing, shell scripting, command line | Computer Science | Command-line tool | Documentation, Uses, and more |
| gnuplot | mcc | N/A | GNUPlot 5.4.10 in a Conda environment | Graphing, Visualization, Data Analysis | Computational Science, Physical Sciences | Plotting & Data Visualization | Documentation, Uses, and more |
| go | ecc | N/A | Go is expressive, concise, clean, and efficient. Its concurrency mechanisms make it easy to write programs that get the most out of multicore and networked machines, while its novel type system enables flexible and modular program construction. Go compiles quickly to machine code yet has the convenience of garbage collection and the power of run-time reflection. It's a fast, statically typed, compiled language that feels like a dynamically typed, interpreted language. Description Source: https://go.dev/doc/ |
Programming Language, Open Source, Concurrency, Systems Programming | Software Engineering, Computer Science | Programming Language | Documentation, Uses, and more |
| go-bootstrap | ecc | N/A | go-bootstrap is a Go programming language template project that provides a starting point for building Go applications with a predefined project structure, configuration, and best practices. | Go, Template, Project Structure, Best Practices | Engineering & Technology | Template Project | Documentation, Uses, and more |
| go-md2man | ecc | N/A | A tool for converting Markdown files to man pages. | Markdown, Documentation, Man Pages | Software Engineering, Other Computer and Information Sciences | Command Line Tool | Documentation, Uses, and more |
| gobject-introspection | mcc | N/A | GObject introspection is a middleware layer between C libraries (using GObject) and language bindings. The C library can be scanned at compile time and generate metadata files, in addition to the actual native C library. Then language bindings can read this metadata and automatically provide bindings to call into the C library. Description Source: https://gi.readthedocs.io/en/latest/ |
Middleware, Language Bindings, Automation, Compatibility | Computer Science, Computer & Information Sciences | Library | Documentation, Uses, and more |
| gocryptfs | ecc | N/A | gocryptfs is a user-space encryption tool for filesystems, designed to provide secure encryption of files and directories. | encryption, filesystems, security, data protection | Cryptography, Computer Science | Encryption Tool | Documentation, Uses, and more |
| gperf | mcc, ecc | N/A | GNU gperf is a perfect hash function generator. For a given list of strings, it produces a hash function and hash table, in form of C or C++ code, for looking up a value depending on the input string. The hash function is perfect, which means that the hash table has no collisions, and the hash table lookup needs a single string comparison only. Description Source: https://www.gnu.org/software/gperf/ |
Perfect Hash Function, Hash Table, Optimization, Gnu | Computer Science, Computer & Information Sciences | Compiler/Generator | Documentation, Uses, and more |
| graphaligner | mcc | 1.0 | GraphAligner is a tool for aligning long sequencing reads to sequence/genome graphs, a key operation in graph-based genome assembly and pangenomics where a reference is represented as a variation graph rather than a single linear sequence. It efficiently maps noisy long reads (Nanopore/PacBio) onto graphs in GFA format using a bit-parallel dynamic-programming approach, producing alignments in GAF (or JSON) format that record the path a read takes through the graph. This is used, for example, to place reads on an assembly graph for error correction, to genotype structural variants against a pangenome graph, or to validate and scaffold graph assemblies. It offers presets that tune parameters for the read type, and version 1.0.20 integrates well with the vg and pangenome toolchains. | Documentation, Uses, and more | |||
| graphviz | lcc | N/A | Graphviz is open source graph visualization software. Graph visualization is a way of representing structural information as diagrams of abstract graphs and networks. It has important applications in networking, bioinformatics, software engineering, database and web design, machine learning, and in visual interfaces for other technical domains. | Graph Visualization, Diagramming, Open Source Software | Graph Theory, Computer & Information Sciences | Data Visualization | Documentation, Uses, and more |
| grep | ecc | N/A | grep is a command-line utility for searching plain-text data sets for lines that match a regular expression. | Text Processing, Command Line Tool, Regular Expressions | Computer Science | Command Line Utility | Documentation, Uses, and more |
| gridss | lcc | 1.0 | GRIDSS (Genomic Rearrangement IDentification Software Suite) is a structural-variant and breakend detection toolkit for short-read whole-genome sequencing, built around a novel genome-wide break-end assembler. Rather than calling variants directly from split/discordant reads, GRIDSS performs positional de novo assembly of reads spanning putative breakpoints, then realigns the assembly contigs to precisely localize breakends and score them, giving high sensitivity and base-pair-resolution calls for deletions, duplications, inversions, insertions and translocations. It consumes coordinate-sorted BAM alignments plus a reference FASTA and emits a VCF describing individual breakends (BND records) that can be paired into structural-variant events. GRIDSS is commonly run through its driver script; this 2.x line is the GRIDSS2 generation, whose output feeds downstream somatic-cancer tools such as GRIPSS (breakend filtering), PURPLE (purity/ploidy and copy number) and LINX (SV clustering and interpretation). This is version 2.13.2 from bioconda. | Structural Variant Calling, DNA Sequencing, Genomic Rearrangements | Genetics, Biological Sciences | Tool | Documentation, Uses, and more |
| gromacs | mcc, lcc | 1.0 | GROMACS is a very fast, widely used molecular-dynamics package for simulating the Newtonian equations of motion of biomolecular systems—proteins, lipids, nucleic acids, and their aqueous/membrane environments—as well as some soft-matter and materials systems. It is renowned for its highly optimized, SIMD- and GPU-accelerated kernels and its efficient domain-decomposition parallelization. The workflow moves through a set of gmx tools: pdb2gmx to build the topology, editconf/solvate/genion to set up the box and solvent, grompp to combine coordinates, topology, and run parameters (.mdp) into a portable binary run input (.tpr), and mdrun to run the simulation, followed by a rich set of analysis tools (RMSD, radius of gyration, hydrogen bonds, free energy, etc.). This CUDA-enabled 2023.3 build targets NVIDIA P100/V100/A100/H200 GPUs and provides both a thread-MPI binary (gmx) and a full-MPI binary (gmx_mpi) with a UCX/PMIx/OpenMPI stack for InfiniBand and SLURM multi-node runs. | Molecular Dynamics, Simulation, High Performance Computing, Biomolecular Systems | Chemistry, Biological Sciences | Molecular Dynamics Software | Documentation, Uses, and more |
| gromacs-plumed2 | lcc | N/A | GROMACS 2019.2 built with the Intel compiler and patched with PLUMED2. GROMACS is a high-performance molecular dynamics engine for simulating biomolecules (proteins, lipids, nucleic acids) and soft-matter systems; the PLUMED2 plug-in adds enhanced-sampling and free-energy methods such as metadynamics and collective-variable-based biasing. It is used in computational chemistry, biophysics, and drug-discovery research. | Molecular Dynamics, Simulation, Biophysics, Computational Chemistry | Biophysics, Biochemistry and Molecular Biology | Simulation Software | Documentation, Uses, and more |
| gromacs-plumed2-gpu | lcc | N/A | GROMACS 2019 (Intel-compiled) with the PLUMED2 plugin and GPU acceleration — a high-performance molecular dynamics engine for simulating biomolecules, polymers, and soft matter. The PLUMED2 integration adds enhanced-sampling and free-energy methods (metadynamics, umbrella sampling, collective-variable analysis). Used across biophysics, biochemistry, and materials modeling. | Documentation, Uses, and more | |||
| grompp | mcc | 1.0 | GROMACS preprocessor from the SMOG structure-based model build of GROMACS 4.5.4, which turns a topology, a coordinate file and a run parameter file into the binary run input consumed by mdrun. This build understands Gaussian contact potentials in the pairs section of the topology, function types 5, 6 and 7, which ordinary GROMACS releases reject outright. Function type 6 combines a Gaussian well with an r^-12 repulsive core and holds its minimum at a chosen distance and depth, which is the form normally wanted for coarse-grained native-contact models of protein folding. Pairs using types 6 or 7 must also be listed as exclusions so the repulsion is not counted twice. Both single and double precision preprocessors are available, the double precision one carrying a _d suffix. | Documentation, Uses, and more | |||
| gsalign | mcc | 1.0 | GSAlign (bioconda 1.0.22) is a fast, memory-efficient sequence aligner designed specifically for comparing pairs of highly similar whole genomes, such as different assemblies or strains of the same or closely related species. It uses an FM-index (BWT-based) seeding strategy to rapidly find maximal exact matches, then extends and clusters them to identify both local and global alignments and to call sequence variants (SNPs and indels) between the two genomes. Because it is optimized for high-identity comparisons, it is substantially faster and lighter on memory than general-purpose whole-genome aligners for that use case. Inputs are a reference genome FASTA (which is first indexed) and a query genome FASTA; outputs include alignment blocks (e.g., MAF/BED-style) and a VCF of detected variants. It is well suited to intraspecific variant discovery and pangenome comparison workflows. | Documentation, Uses, and more | |||
| gsl | mcc, lcc | N/A | GNU Scientific Library (GSL) | Numerical Library, Mathematical Functions, C Programming, C++ Programming | Mathematics, Other Mathematics | Library | Documentation, Uses, and more |
| gtdbtk | mcc | 1.0 | GTDB-Tk is a toolkit for assigning objective, standardized taxonomic classifications to bacterial and archaeal genomes based on the Genome Taxonomy Database (GTDB), a phylogenetically consistent, rank-normalized taxonomy. Its main classify_wf workflow identifies 120 bacterial (or 53 archaeal) single-copy marker genes with HMMER and Prodigal, aligns them, places each query genome into a reference tree with pplacer, and refines classification using average nucleotide identity (ANI) via FastANI and relative evolutionary divergence. It takes as input a directory of genome FASTA files (isolates or MAGs) and produces summary TSV files giving the full GTDB taxonomic lineage, closest reference genome, ANI, and any classification caveats, and it requires a large downloaded GTDB reference release. This is bioconda gtdbtk 2.3.2. | bioinformatics, genomics, taxonomy, data analysis | Bioinformatics, Biological Sciences | Command-line tool | Documentation, Uses, and more |
| gtf2csv | mcc, lcc | 1.0 | gtf2csv is a small Python command-line utility (from github.com/zyxue/gtf2csv) that converts GTF genome-annotation files into flat, tabular CSV (or pickle) form so the many key=value attributes packed into the GTF 9th column are expanded into individual, easily queryable columns. Standard GTF is awkward to parse for ad-hoc analysis because feature attributes are embedded in a semicolon-delimited string; gtf2csv flattens these into a proper table (e.g. one column each for gene_id, gene_name, gene_biotype, transcript_id, etc.), which loads directly into pandas, R, or a spreadsheet. It converts an annotation GTF into a corresponding .csv, with options for output format and multiprocessing. This is a lightweight preprocessing convenience rather than an analysis tool, valuable whenever a workflow needs to filter, join, or summarize annotation records programmatically. | Documentation, Uses, and more | |||
| gtk-doc | mcc | N/A | gtk-doc is a documentation generation tool for GTK and GNOME projects. It allows developers to create API documentation from comments in the source code. | Documentation, API, GTK, GNOME | Software Engineering, Computer Science | Documentation Generator | Documentation, Uses, and more |
| gtkplus | mcc | N/A | GTK+ is a multi-platform toolkit for creating graphical user interfaces. It is used by many applications to provide a consistent look and feel across different operating systems. | GUI, Toolkit, Cross-platform, Open Source | Software Engineering, Computer Science | Library | Documentation, Uses, and more |
| gubbins | lcc | 1.0 | Gubbins (Genealogies Unbiased By recomBinations In Nucleotide Sequences) is a tool for building whole-genome phylogenies from closely related bacterial isolates while accounting for the confounding effect of homologous recombination. Because recombined regions carry an atypical density of base substitutions that would otherwise distort a clonal tree, Gubbins iteratively identifies genomic regions with elevated SNP density (likely recombination), removes them from consideration, and rebuilds the phylogeny from the remaining vertically inherited (clonal) variation, converging on a recombination-corrected tree. Its input is a whole-genome multi-FASTA alignment of the isolates (typically produced by mapping to a common reference and calling a core-SNP or full alignment), and it can optionally choose the tree-builder (RAxML, FastTree, IQ-TREE) and substitution model. Outputs include the final recombination-filtered phylogeny (Newick), a GFF of predicted recombination events per branch, and per-branch substitution statistics, which are commonly visualized against the tree (e.g. with Phandango). This is version 3.4 from Bioconda. | Bacterial Recombination, Whole-Genome Sequencing, Population Genetics | Bioinformatics, Biological Sciences | Standalone | Documentation, Uses, and more |
| guidance | lcc | 1.0 | GUIDANCE (version 2.02) assesses the reliability of multiple sequence alignments (MSAs) and flags the columns and sequences whose alignment is uncertain, so that unreliable regions can be masked before phylogenetic or selection analyses that are sensitive to alignment error. It estimates confidence by perturbing the guide tree and/or bootstrapping and realigning, then measuring how consistently each residue pair, column, and sequence is aligned across the alternative alignments, yielding per-column and per-sequence confidence scores. It supports several alignment programs as back-ends (MAFFT, PRANK, MUSCLE, CLUSTALW) and offers the GUIDANCE, HoT, and GUIDANCE2 algorithms. A run takes unaligned sequences and a chosen aligner, and outputs the MSA, color-coded reliability scores, and filtered alignments with low-confidence columns/sequences removed at a user-set threshold. It is commonly used to improve robustness of downstream tree inference and dN/dS tests. | Documentation, Uses, and more | |||
| guppy | mcc | 1.0 | Guppy is Oxford Nanopore Technologies' production basecalling and demultiplexing software, converting the raw electrical-signal data (fast5) produced by nanopore sequencers into base-called nucleotide reads in FASTQ/FASTA. It applies neural-network basecalling models (fast/high-accuracy configurations) and additionally supports adapter trimming, barcode demultiplexing, and quality filtering of the resulting reads. It is run on the command line via a basecaller executable that takes a fast5 directory and a config appropriate to the flowcell/kit, with a companion barcoder utility for demultiplexing. This is the CPU build of version 3.0.3 (an older release; note it predates the fast5-to-pod5 transition and the later Dorado basecaller), suitable where GPU acceleration is unavailable. | Bioinformatics, Sequencing, Nanopore Sequencing, Genomics, DNA Sequencing | Genomics, Biological Sciences | Basecaller | Documentation, Uses, and more |
| gyr-curation-asm-path-translate | mcc | 1.0 | This is a small helper utility (asm-path-translate-printout.py) from the Gyr Genome Assembly Curation repository, used during manual curation of long-read genome assemblies produced by graph-based assemblers such as Verkko and hifiasm. Those assemblers describe contigs/scaffolds as paths through a unitig graph using notation like utig4-0000001, and this script translates or prints out such assembly-path notation (e.g. mapping utig4 to utig1 level identifiers) so that a curator can interpret and follow the graph structure while editing. It is a command-line Python script operating on assembly path/notation files, serving as a convenience during the interactive curation of difficult genome regions rather than as a standalone assembler. It is provided alongside Verkko and related long-read curation tools in the container. | Documentation, Uses, and more | |||
| gyr-curation-patch2path | mcc | 1.0 | This is the patch_2_path.sh driver script from the Gyr Genome Assembly Curation helper repository (SEFumagalli/Gyr-Genome-Assembly-Curation), a set of manual-curation utilities developed around assembling the Gyr (Bos indicus) genome. patch_2_path automates part of the gap-patching / assembly-path editing step for graph-based long-read assemblies produced by Verkko or hifiasm, translating curated 'patches' (corrections that resolve gaps or fix mis-joins) into updated assembly paths through the sequence graph. It is a workflow glue script rather than a general-purpose tool, intended to be run within a manual-curation loop where a curator inspects the assembly graph, identifies problem regions, and applies path edits to improve contiguity and correctness. It fits into the broader Gyr curation toolkit alongside Verkko, alignment, and evaluation steps, and is packaged here so that path-patching can be reproduced on the cluster. | Documentation, Uses, and more | |||
| harfbuzz | mcc | N/A | HarfBuzz is a text shaping library. Using the HarfBuzz library allows programs to convert a sequence of Unicode input into properly formatted and positioned glyph output—for any writing system and language. Description Source: https://harfbuzz.github.io/ |
Text Shaping Engine, Unicode Text, Text Layout, Font Features Customization | Software Engineering, Computer & Information Sciences | Library | Documentation, Uses, and more |
| hdf5 | mcc, lcc | N/A | A general purpose library and file format for storing scientific data | File Format, Data Management, Data Storage, Data Sharing, High-Performance Computing | Computer Science | Library/Tool | Documentation, Uses, and more |
| help2man | mcc, lcc | N/A | help2man is a tool for automatically generating simple manual pages from program output. Description Source: https://www.gnu.org/software/help2man/#Overview |
Documentation, Man Pages, Automation | Software Engineering, Systems & Development, Engineering & Technology | Utility | Documentation, Uses, and more |
| herro | lcc | 1.0 | HERRO is a deep-learning-based error-correction tool for Oxford Nanopore long reads that is haplotype-aware, meaning it corrects reads while preserving true heterozygous differences rather than collapsing haplotypes. It uses all-vs-all read overlaps (via minimap2) to build alignment features and a neural network (running on libtorch/PyTorch with GPU acceleration) to infer the corrected sequence, producing highly accurate long reads that substantially improve the contiguity and correctness of downstream de novo assemblies. The workflow typically involves a preprocessing/alignment step to generate read overlaps followed by the inference step that emits corrected reads in FASTA/FASTQ. This build compiles HERRO from the dominikstanojevic GitHub repository with cargo (Rust), within a conda environment providing its dependencies, alongside minimap2 and CUDA 11.7 libtorch, and runs on GPU (CUDA 11.7). | Documentation, Uses, and more | |||
| hic-pro | lcc | N/A | Hic-Pro | Bioinformatics, Genomics, Chromatin Structure, Hi-C Data, Data Visualization, Chromatin Interactions | Genomics, Biological Sciences | Data Analysis | Documentation, Uses, and more |
| hicpro | mcc | 1.0 | HiC-Pro is an established, optimized pipeline for processing Hi-C chromosome-conformation-capture sequencing data from raw FASTQ reads all the way to normalized interaction/contact maps. It performs independent (split) mapping of the two read ends with Bowtie2, then rescues chimeric reads spanning ligation junctions, assigns aligned reads to restriction fragments, filters out invalid ligation products (dangling ends, self-circles, re-ligation, and duplicates), and builds genome-wide contact matrices at multiple resolutions, which it then normalizes using the iterative correction (ICE) method. It supports both restriction-enzyme-based Hi-C and enzyme-free protocols, allele-specific analysis with phased SNPs, and can process large datasets in parallel on a cluster. Configuration is via a single config-hicpro.txt file specifying genome index, restriction fragments, and resolutions. Outputs (validPairs and .matrix/.bed files) can be converted for visualization in Juicebox or cooler/HiGlass. This deployment builds HiC-Pro from the nservant/HiC-Pro GitHub repo via its environment.yml and make install. | bioinformatics, genomics, data analysis | Computational Biology, Bioinformatics | Analysis Tool | Documentation, Uses, and more |
| hicup | lcc | N/A | Documentation, Uses, and more | ||||
| hifiasm | mcc | 1.0 | hifiasm (bioconda 0.24.0) is a fast, haplotype-resolved de novo assembler purpose-built for PacBio HiFi (circular consensus) reads. It uses a phased string-graph model to preserve haplotype information, producing not only a primary contig set but also two partially or fully phased haplotype assemblies. Beyond HiFi-only assembly, it supports Hi-C-integrated phasing (with paired Hi-C reads) and trio binning (using parental k-mer databases from yak) to achieve chromosome-scale haplotype separation, and it can incorporate ultra-long Oxford Nanopore reads to improve contiguity. Inputs are HiFi FASTQ/FASTA reads (optionally plus Hi-C and/or parental read sets); outputs are GFA graph files that are converted to FASTA contigs, including hap1/hap2 assemblies. It is widely regarded as a state-of-the-art assembler and is a standard input to purge_dups and Hi-C scaffolding. | Assembler, De Novo Assembly | Biological Sciences | Bioinformatics | Documentation, Uses, and more |
| hifiasm-0.25.0 | mcc | 1.0 | hifiasm (version 0.25.0) is a fast, haplotype-resolved de novo assembler built primarily for PacBio HiFi (high-fidelity) long reads. It constructs a phased string/assembly graph and can produce partially or fully phased contigs; with additional data it resolves haplotypes even more completely — accepting parental short reads for trio binning, Hi-C reads for Hi-C-based phasing, and (in recent versions) ultra-long ONT reads to improve contiguity. Outputs are primary and alternate assemblies or a pair of fully phased haplotype assemblies, written as GFA graphs (convertible to FASTA). hifiasm is prized for its speed, low error rate, and ability to generate high-quality, contiguous genomes with minimal parameter tuning. It supports HiFi-only assembly as well as Hi-C-integrated phasing (supplying paired Hi-C read files) and trio mode (using parental k-mer databases built with yak). It is a leading assembler for reference-grade and pangenome projects. | genome assembly, bioinformatics, long-read sequencing | Computational Biology, Bioinformatics | Assembly Tool | Documentation, Uses, and more |
| hisat | lcc | 1.0 | HISAT2 is a fast and memory-efficient spliced aligner for mapping next-generation sequencing reads to a reference genome, most commonly used for RNA-seq where reads span exon-exon junctions. It uses a graph FM-index (a hierarchical set of a whole-genome index plus many small local indexes) that lets it align spliced reads accurately while keeping the index small enough to run on modest memory, and it can incorporate known splice sites and SNPs. The workflow first builds an index from the reference and then aligns reads against it, producing SAM output with spliced-alignment (XS strand) tags suitable for downstream transcript assembly with StringTie. Version 2.2.1 is a standard aligner in RNA-seq quantification pipelines and can also perform DNA and graph-based alignment. | alignment, genomics, bioinformatics | Bioinformatics, Genomics | Command-line tool | Documentation, Uses, and more |
| hisat2 | mcc, lcc | 1.0 | HISAT2 is a fast, low-memory, splice-aware read aligner used chiefly to map RNA-seq reads (and also DNA reads) to a reference genome, correctly spanning exon-exon junctions in spliced transcripts. It achieves both speed and small memory footprint through a hierarchical graph FM-index scheme that combines one global genome index with many small local indexes, and it can incorporate known splice sites and genetic variants (a graph/HISAT-genotype capability) into the index. The workflow first builds a genome index with the hisat2-build tool (optionally supplying splice-site and exon files extracted from a GTF) and then aligns paired or single reads against that index, reporting spliced alignments and a summary of alignment/junction statistics suitable for downstream transcript assembly (StringTie) or quantification. Version 2.2.1 is provided as its own conda environment/app within a large multi-tool Miniconda base container that also bundles phylogenetics, alignment, assembly, and variant-analysis tools. | Alignment, Ngs, Genomics, Transcriptomics | Biology, Biological Sciences | Alignment Tool | Documentation, Uses, and more |
| hmmer | mcc, lcc | 1.0 | HMMER is a widely used sequence-homology search package that detects remote evolutionary relationships more sensitively than simple pairwise aligners by representing families as profile hidden Markov models that capture position-specific residue preferences and insertion/deletion probabilities. Its programs include hmmbuild (build a profile from a multiple alignment), hmmsearch (search a profile against a sequence database), hmmscan (search sequences against a profile database such as Pfam), phmmer and jackhmmer (iterative single-sequence searches), and nhmmer for DNA/nucleotide profile searches. Version 3.x uses a fast heuristic acceleration pipeline that makes profile searches competitive in speed with BLAST while reporting rigorous E-values and bit scores, with per-domain and per-sequence hit tables. A typical use is annotating proteins by running hmmscan against a pressed Pfam database. Version 3.4 is provided as a conda tool within a broad Rocky 8 microbiome/genomics/ML container. | Bioinformatics, Computational Biology, Proteomics, Sequence Analysis | Bioinformatics, Biological Sciences | Bioinformatics Tool | Documentation, Uses, and more |
| hmmer2go | lcc | 1.0 | HMMER2GO is a Perl-based bioinformatics pipeline for functional annotation of protein-coding sequences with Gene Ontology (GO) terms. It first identifies putative open reading frames (ORFs) within nucleotide sequences (e.g., transcriptome assemblies or genomic contigs) via EMBOSS getorf, translates them, and then searches the predicted proteins against the Pfam-A database using HMMER's hmmscan to detect conserved protein domains. Pfam domain hits are mapped to GO terms via the pfam2go mapping, and results can be summarized as GO annotations suitable for downstream enrichment analysis. The tool is organized as a set of subcommands: getorf extracts and translates ORFs, run performs the hmmscan search against Pfam, mapterms assigns GO terms from the Pfam hits, and map2gaf produces a GAF-style association file. It is well suited for annotating non-model-organism transcriptomes where curated annotations are unavailable. | Documentation, Uses, and more | |||
| hocort | mcc | 1.0 | HoCoRT (Host Contamination Removal Tool, bioconda 1.2.2) is a Python-based utility for removing host-derived reads from sequencing datasets, a critical cleaning step in metagenomic, microbiome, and pathogen-detection studies where samples are dominated by host DNA/RNA. Rather than committing to a single method, it provides a unified interface to a choice of underlying read mappers and classifiers — including Bowtie2, HISAT2, BWA-MEM2, Minimap2, Kraken2, and combinations thereof — letting the user trade off sensitivity, specificity, and speed appropriate to short or long reads. Reads mapping to the supplied host reference (or classified as host) are filtered out, and the cleaned, non-host reads are written for downstream analysis. Inputs are FASTQ reads (single- or paired-end) plus a prebuilt host index. HoCoRT also provides an index-building subcommand for each supported mapper. | Documentation, Uses, and more | |||
| homer | mcc, lcc | 1.0 | HOMER (Hypergeometric Optimization of Motif EnRichment) is a suite of command-line tools for motif discovery and next-generation sequencing analysis, originally developed for ChIP-seq but broadly applied to ATAC-seq, RNA-seq, DNase-seq, Hi-C, and GRO-seq. Its signature capability is de novo and known motif enrichment analysis: findMotifsGenome.pl and findMotifs.pl scan a set of genomic regions or gene promoters against a background and report enriched transcription-factor binding motifs with enrichment p-values. HOMER also provides peak/transcript finding (findPeaks), a flexible tag-directory model (makeTagDirectory) that pre-processes aligned reads into a compact form, coverage/BedGraph generation (makeUCSCfile), quantification and annotation of peaks relative to genes (annotatePeaks.pl), and differential and metagene analyses. A typical workflow builds a tag directory from aligned reads, calls peaks, then runs motif enrichment on the peaks against a specified genome to identify candidate regulators. Note that HOMER requires downloading genome and promoter data packages separately via its configureHomer.pl script. This is version 4.11 from Bioconda. | Bioinformatics, Next-Generation Sequencing, Motif Discovery, Transcription Factor Binding Sites | Genetics, Biological Sciences | Tool | Documentation, Uses, and more |
| htseq | lcc | 1.0 | HTSeq is a Python library and set of command-line scripts for processing high-throughput sequencing data, best known for its htseq-count utility for gene-level read quantification in RNA-seq differential-expression workflows. htseq-count takes aligned reads (a SAM/BAM file, name- or position-sorted) together with a genome annotation (GTF/GFF) and counts how many reads overlap each feature (typically each gene), applying a user-selected overlap-resolution mode (union, intersection-strict, or intersection-nonempty) and strandedness setting (yes/no/reverse), and reporting special categories such as no_feature, ambiguous, and alignment_not_unique. The resulting per-gene count table feeds directly into DESeq2, edgeR, or limma. Beyond the counting script, the HTSeq Python API provides classes for parsing FASTQ/SAM/BAM/GFF/VCF and for building genomic-interval arrays, enabling custom coverage and annotation analyses. This build is version 0.12.4 from Bioconda. | Bioinformatics, RNA-Seq, High-Throughput Sequencing, Gene Expression, Sequence Data | Bioinformatics, Biological Sciences | Bioinformatics Tool | Documentation, Uses, and more |
| htslib | mcc | N/A | HTSlib is an implementation of a unified C library for accessing common file formats, such as SAM, CRAM and VCF, used for high-throughput sequencing data, and is the core library used by samtools and bcftools. HTSlib only depends on zlib. It is known to be compatible with gcc, g++ and clang. Description Source: https://github.com/samtools/htslib |
Bioinformatics, Sequencing Data, File Format Processing, C Library | Genomics, Biological Sciences | Data Processing | Documentation, Uses, and more |
| humann | mcc, lcc | 1.0 | HUMAnN 3 (version 3.9), part of the bioBakery suite, profiles the functional potential of microbial communities from shotgun metagenomic or metatranscriptomic sequencing — quantifying the presence and abundance of metabolic pathways and gene families rather than just taxonomy. It uses a tiered search strategy: MetaPhlAn identifies which species are present, reads are mapped against pangenomes of those species (via Bowtie2) for known organisms, and unmapped reads undergo a translated search (DIAMOND) against a protein database (UniRef90/UniRef50). Outputs are three tables — gene-family abundances, pathway abundances, and pathway coverage — reported in RPK and stratified by contributing species, which can be normalized (e.g. to CPM/relative abundance) and regrouped to KO, EC, GO, or other ontologies with utility scripts. It is a cornerstone of functional metagenomics and is often paired with MetaPhlAn in the bioBakery workflows. | Metagenomics, Microbiome, Pathway Analysis, Bioinformatics | Biological Sciences | Analysis Tool | Documentation, Uses, and more |
| humann2 | lcc | 1.0 | HUMAnN2 (HMP Unified Metabolic Analysis Network, version 2, biobakery build 2.8.1) is a pipeline for functional profiling of shotgun metagenomic and metatranscriptomic sequencing data, quantifying the presence and abundance of microbial metabolic pathways, gene families, and enzymes and attributing them to contributing species. It uses a tiered search strategy: it first identifies which species are present with MetaPhlAn, builds a sample-specific pangenome database and maps reads to it with a nucleotide aligner (Bowtie2), then translates and searches the remaining unmapped reads against a protein database (DIAMOND) for broader functional coverage. Results are reported as gene-family abundances (UniRef), pathway abundance, and pathway coverage, with a stratified breakdown showing which community members contribute each function, and companion utilities can renormalize (to CPM/relative abundance) and regroup gene families into other ontologies (KO, EC, GO, etc.). Inputs are quality-controlled metagenomic reads (FASTQ), optionally with a precomputed taxonomic profile. It is a core biobakery tool for community metabolic-potential analysis (superseded by HUMAnN 3 but still used). | Genomics | Documentation, Uses, and more | ||
| hwloc | mcc, lcc, ecc | N/A | Portable Hardware Locality | Machine Topology, Hardware Architecture, System Software | Computer Science, Computer & Information Sciences | Library | Documentation, Uses, and more |
| hyphy | lcc | 1.0 | HyPhy (Hypothesis testing using Phylogenies, davebx build 2.5.17) is a software package and scripting environment for maximum-likelihood analysis of molecular sequence evolution, focused especially on detecting and characterizing natural selection from codon-aware alignments and phylogenetic trees. It implements a well-known suite of selection-analysis methods, including FEL and SLAC (site-level dN/dS), MEME (episodic diversifying selection at sites), FUBAR (Bayesian site-level selection), aBSREL and BUSTED (branch/gene-level episodic selection), RELAX (tests for relaxation/intensification of selection), and GARD (recombination breakpoint detection). Inputs are a multiple-sequence alignment (typically in-frame codon nucleotide FASTA/NEXUS) and a phylogenetic tree; outputs are JSON result files plus human-readable summaries reporting per-site or per-branch selection signals. It can be run interactively, through named analyses such as MEME, or via its batch language (HBL) for custom models. It underpins the Datamonkey web server and is a standard tool in molecular-evolution studies. | Bioinformatics, Molecular Evolution, Phylogenetics, Epidemiology | Bioinformatics, Biological Sciences | Open-Source Software | Documentation, Uses, and more |
| hyphy-analysis | lcc | N/A | HyPhy-analyses is a collection of standardized phylogenetic and molecular-evolution analysis workflows built on the HyPhy (Hypothesis testing using Phylogenies) platform (github.com/veg/hyphy-analyses). It is used in evolutionary biology to detect selection, recombination, and other sequence-evolution signals from alignments and phylogenies. | bioinformatics, phylogenetics, molecular evolution, statistical analysis | Computational Biology, Molecular Evolution | Analysis Tool | Documentation, Uses, and more |
| hypo | mcc | 1.0 | HyPo (Hybrid Polisher) is a fast genome-assembly polishing tool that corrects base-level errors (mismatches and small indels) in draft long-read assemblies by leveraging both short and long reads. It exploits highly accurate short reads for the primary error signal while using long reads to resolve context in repetitive regions, and it is engineered to be substantially faster and more memory-efficient than comparable polishers, making it practical for large genomes and multiple polishing rounds. It takes the draft assembly FASTA together with alignments of short reads (and optionally long reads) to that draft, plus estimates of coverage and genome size, and outputs a polished, error-reduced assembly. This is version 1.0.3 from bioconda. | Statistical Analysis, Hypothesis Testing, Data Analysis | Statistics & Probability | Scientific Software | Documentation, Uses, and more |
| hypre | mcc, lcc | N/A | Scalable algorithms for solving linear systems of equations | Preconditioners, Solvers, Linear Systems, Sparse Matrices | Applied Mathematics, Mathematics | Computational Software | Documentation, Uses, and more |
| icc | lcc | N/A | Intel C++ Compiler (icc) is a high-performance compiler for C and C++ programming languages, optimized for Intel architectures. | compiler, C++, performance, parallel computing | Software Engineering, Computer Science | Compiler | Documentation, Uses, and more |
| icc32 | lcc | N/A | Documentation, Uses, and more | ||||
| icu4c | mcc | N/A | ICU (International Components for Unicode) is a mature, widely used set of C/C++ and Java libraries providing Unicode and Globalization support for software applications. ICU is widely portable and gives applications the same results on all platforms and between C/C++ and Java software. | Unicode, Globalization, Cross-Platform, Libraries | Software Engineering, Computer & Information Sciences | Development | Documentation, Uses, and more |
| idr | mcc, lcc | 1.0 | IDR (Irreproducible Discovery Rate) is a statistical method and command-line tool from the ENCODE project used in functional genomics to assess the reproducibility of high-throughput ChIP-seq, ATAC-seq, and similar peak-calling experiments across biological replicates. It models the ranked signal of peaks from two replicates as a mixture of a reproducible (consistent) component and an irreproducible (noise) component, fitting a copula-based model to assign each peak an IDR value. Peaks are supplied as narrowPeak/broadPeak/BED-like files ranked by a chosen score column (e.g. signal value, p-value, or q-value), and the tool outputs a merged peak set annotated with local and global IDR scores plus optional plots of the fit. Peaks below an IDR threshold such as 0.05 are retained to define a high-confidence, reproducible peak set. This build is bioconda idr 2.0.4.2. | Documentation, Uses, and more | |||
| ifort | lcc | N/A | Intel® Fortran Compiler Classic for building optimized Intel® 64, and IA-32 CPUs | Fortran, Compiler, High Performance Computing, Scientific Computing | Computer Science, Software Engineering | Compiler | Documentation, Uses, and more |
| ifort32 | lcc | N/A | Intel® Fortran Compiler Classic for building optimized Intel® 64, and IA-32 CPUs | Documentation, Uses, and more | |||
| igv | mcc | N/A | IGV software. | Visualization, Genomics, Sequencing, Data Analysis | Bioinformatics, Biological Sciences | Genomics Software | Documentation, Uses, and more |
| image-video-studio | lcc, ecc | 1.0 | Image-Video-Studio is an Open OnDemand interactive application registration that gives ECC cluster users a browser-based workspace for image and video processing and editing, launched as an OOD interactive session (batch-connect job) that runs on a compute node and is reached through the OnDemand web portal. Rather than a single command-line binary, it provides a graphical environment within the browser for common media tasks, letting researchers work with images and video without configuring X forwarding or local software. In the SDS catalog this entry is a catalog-only stub: it is not a built Singularity image but a registration that surfaces the ECC cluster's OnDemand interactive apps so users can discover them. Users typically start it from the OnDemand Interactive Apps menu, choosing resource parameters (cores, memory, time), then connect to the running session in their browser. | Documentation, Uses, and more | |||
| imagemagick | mcc, lcc | 1.0 | ImageMagick https://imagemagick.org/script/install-source.php | Image Processing, Graphics Editing, Conversion Tool | Image Processing, Computer & Information Sciences | Utility | Documentation, Uses, and more |
| imb | mcc, lcc | N/A | Intel MPI Benchmarks (IMB) | Benchmarking, In-Memory Databases | Computer & Information Sciences | Tool | Documentation, Uses, and more |
| imlib2 | mcc | N/A | Imlib2 is an image loading and rendering library that provides a simple API for handling images in various formats. | Image Processing, Graphics, Library | Image Processing, Computer Science | Library | Documentation, Uses, and more |
| impi | lcc | N/A | Intel MPI Library (C/C++/Fortran for x86_64) | Message Passing Library, Parallel Computing, Asynchronous Messaging | High-Performance Computing, Engineering & Technology | Communication Library | Documentation, Uses, and more |
| impi/2018.3.222 | lcc | N/A | Intel MPI Library (C/C++/Fortran for x86_64) | Documentation, Uses, and more | |||
| impi/2019.3.199 | lcc | N/A | Intel MPI Library (C/C++/Fortran for x86_64) | Documentation, Uses, and more | |||
| impi/2019.4.243 | lcc | N/A | Intel MPI Library (C/C++/Fortran for x86_64) | Documentation, Uses, and more | |||
| implicit-landuse | lcc | 1.0 | Documentation, Uses, and more | ||||
| init_opencl | mcc, lcc | N/A | init_opencl is a software library that is used to initialize OpenCL platforms and create OpenCL contexts for parallel computing applications. | Software, Compiler, Library, Hpc | Computer & Information Sciences | Software Development | Documentation, Uses, and more |
| inotify-tools | mcc, lcc | 1.0 | inotify-tools is a small package of command-line utilities that expose the Linux kernel's inotify filesystem-event notification API, letting scripts react to file and directory changes without polling. It provides two programs: inotifywait, which blocks until one or more specified filesystem events (such as create, modify, delete, move, close_write, or access) occur on watched files/directories and then prints or exits — ideal for triggering an action whenever a file appears or changes; and inotifywatch, which collects and tabulates event statistics over a period to show how frequently different events occur. Common patterns include a shell loop that waits for a close_write event and then runs a processing step, and recursive, continuous monitoring of a directory tree for an event stream. It is frequently used to build lightweight file watchers, auto-rebuild/auto-sync triggers, and pipeline hand-off mechanisms. This build is version 3.20.2 from conda-forge; note it relies on Linux inotify and therefore watches local filesystems (not all network mounts emit events). | Linux, File Monitoring, System Tools | Systems, Other Computer and Information Sciences | Command-line Tools | Documentation, Uses, and more |
| inputproto | mcc, lcc | N/A | A package that provides the headers used to compile Xlib clients. | Software Development, Headers, Xlib | Computer & Information Sciences | Library | Documentation, Uses, and more |
| inspector | mcc, lcc | 1.0 | Inspector is a reference-free tool for evaluating the accuracy of long-read de novo genome assemblies by mapping the original long reads back to the assembly and analyzing the alignments to detect both large structural errors (expansions, collapses, inversions, haplotype switches) and small-scale errors (base substitutions and short indels). Beyond reporting these errors and standard contiguity statistics, it can also correct them, producing a polished assembly, via its correction step. The workflow first evaluates the assembly against the reads (using minimap2 for alignment and pysam/statsmodels/flye among its dependencies) for HiFi, Nanopore, or CLR data, yielding a summary of assembly QV, error counts, and locations, then optionally applies fixes in the correction step. Because it needs no reference genome, it is well suited to non-model-organism assemblies where only the sequencing reads and the draft assembly are available. | Monitoring, Security, Network Analysis | Computer Science, Computer & Information Sciences | Security Software | Documentation, Uses, and more |
| intel | lcc | N/A | Intel Compiler Family (C/C++/Fortran for x86_64) | Compiler, Optimization, Performance, Debugging | Computer and Information Sciences | Development Tools | Documentation, Uses, and more |
| intel-oneapi-compilers | mcc, lcc | N/A | Intel oneAPI Compilers provide a unified programming model for high-performance computing, enabling developers to optimize applications across various architectures including CPUs and GPUs. | compilers, high-performance computing, Intel, oneAPI, C++, Fortran | Computer Science, Software Engineering | Compiler | Documentation, Uses, and more |
| intel-oneapi-mkl | mcc | N/A | Intel oneAPI Math Kernel Library (MKL) provides highly optimized mathematical functions for scientific computing, including linear algebra, fast Fourier transforms, and vector mathematics. | Mathematics, High-Performance Computing, Numerical Libraries | Numerical Analysis, Applied Mathematics | Library | Documentation, Uses, and more |
| intel-oneapi-mpi | mcc, lcc | N/A | Intel oneAPI MPI is a high-performance message passing interface library designed to facilitate parallel programming in distributed computing environments. | MPI, Parallel Computing, High Performance Computing, Distributed Systems | Computer Science, Computer and Information Sciences | Library | Documentation, Uses, and more |
| intel-oneapi-tbb | mcc | N/A | Intel oneAPI Threading Building Blocks (TBB) is a C++ template library for parallel programming that abstracts low-level threading details, allowing developers to write scalable and high-performance applications. | Parallel Computing, C++, Performance, Multithreading | Computer Science, Software Engineering | Library | Documentation, Uses, and more |
| intel-parallel-studio-cluster.2019.5-gcc | mcc | N/A | Intel Parallel Studio Cluster is a suite of tools for optimizing and debugging high-performance applications, particularly for clusters and cloud environments. | HPC, Parallel Computing, Performance Optimization, Debugging Tools | Computer Science | Development Tools | Documentation, Uses, and more |
| intel-parallel-studio-cluster.2020.4-gcc | mcc | N/A | Intel Parallel Studio Cluster is a suite of tools for developing high-performance applications on clusters, providing advanced capabilities for parallel programming, performance analysis, and debugging. | HPC, Parallel Computing, Performance Analysis, Debugging, Cluster Computing | Computer Science | Development Tools | Documentation, Uses, and more |
| intel_ipp_ia32 | mcc, lcc | N/A | Intel Integrated Performance Primitives (IPP) for IA-32 architecture provides a set of highly optimized software functions for multimedia processing, data processing, and communications. | HPC, Performance Optimization, Multimedia Processing | Computer Science, Computer and Information Sciences | Library | Documentation, Uses, and more |
| intel_ipp_intel64 | mcc, lcc, ecc | N/A | A library of multimedia and data processing optimized for Single Instruction, Multiple Data (SIMD) instructions. | Software Development, Performance Optimization, Data Processing, Signal Processing, Image Processing, Cryptography | Software Engineering, Computer & Information Sciences, Systems & Development | Commercial | Documentation, Uses, and more |
| intel_ippcp_ia32 | mcc, lcc | N/A | Intel Integrated Performance Primitives - Cryptography (IPP-Crypto) is a library that provides highly optimized building blocks for a variety of encryption and decryption algorithms on Intel architecture processors. It aims to accelerate cryptographic operations to enhance performance in software applications. | Cryptographic Library, Performance Optimization, Intel Architecture | Computer & Information Sciences | Compiler/Library | Documentation, Uses, and more |
| intel_ippcp_intel64 | mcc, lcc, ecc | N/A | Intel Integrated Performance Primitives Cryptography (IPP Cryptography) is a collection of highly optimized cryptographic functions and algorithms developed by Intel for high-performance computing applications. It provides a set of cryptographic functions optimized for Intel processors to enhance the security and performance of cryptographic operations in software. | Cryptography, High-Performance Computing, Intel Processors, Security | Computer & Information Sciences | Optimization Library | Documentation, Uses, and more |
| intelpython3_full | lcc | N/A | Intel Distribution for Python is a distribution of Python optimized for performance on Intel architecture, providing libraries and tools for high-performance computing and data analytics. | Python, Data Science, Machine Learning, High Performance Computing, Intel | Data Analysis, Machine Learning | Programming Language Distribution | Documentation, Uses, and more |
| interproscan | mcc, lcc | 1.0 | InterProScan (EBI distribution 5.73-104.0) is the standard tool for large-scale functional annotation of proteins, scanning sequences against the InterPro consortium's member databases — including Pfam, PANTHER, PRINTS, ProSite, CDD, SMART, TIGRFAMs, SUPERFAMILY, Gene3D, and others — to assign protein families, domains, sites, and integrated InterPro entries. From those matches it derives cross-references to Gene Ontology (GO) terms and pathway databases (Reactome, MetaCyc, KEGG mappings), giving a comprehensive functional profile per protein. It runs the many underlying analysis applications in parallel and can also translate and scan nucleotide input via an ORF-finding step. Inputs are protein (or nucleotide) FASTA; outputs are available in TSV, GFF3, XML, and JSON, plus optional GO and pathway annotations. It requires a Java 11 runtime (provided here via the conda environment) and benefits from pre-downloaded member-database data; it is a core step in genome-annotation and comparative-genomics pipelines. | Bioinformatics, Proteomics, Sequence Analysis, Functional Annotation | Bioinformatics, Biological Sciences | Bioinformatics Tool | Documentation, Uses, and more |
| intltool | mcc | N/A | intltool is a set of tools to centralize translation of many different file formats using GNU gettext-compatible PO files. Description Source: https://freedesktop.org/wiki/Software/intltool/ |
Localization, Translation, Internationalization | Computer Science, Computer & Information Sciences | Localization Tool | Documentation, Uses, and more |
| iphop | mcc | 1.0 | iPHoP (Integrating Phage HOst Prediction), version 1.3.3, predicts the host of phage and other viral sequences at the genus level directly from viral genome sequences, without requiring cultivation. Rather than relying on a single signal, it integrates multiple complementary host-prediction methods — including sequence homology to host genomes, CRISPR spacer matches, and machine-learning classifiers trained on genomic features — and combines their scores into a single, confidence-calibrated prediction with an associated score threshold, improving both accuracy and the fraction of viruses that receive a confident host assignment. Input is a FASTA of viral contigs/genomes (typically from tools like VirSorter2, geNomad, or metagenome assembly), and output is a table of predicted host taxa with confidence scores. It depends on a large downloadable reference database of host genomes and markers. It is widely used in viral ecology and metagenomic virome studies to link viruses to their microbial hosts. | Documentation, Uses, and more | |||
| ipyrad | lcc | N/A | ipyrad is an interactive, Python-based toolkit for the assembly and analysis of restriction-site-associated DNA sequencing data (RAD-seq, ddRAD, GBS and related reduced-representation methods). It performs de novo and reference-based read assembly, clustering, filtering, and outputs matrices/format files for downstream phylogenetic and population-genetic analyses. On LCC it is provided as a Miniconda3 conda environment (ipyrad-0.7.30) exposing the ipyrad command-line program. It is a mainstay of evolutionary biology and population genomics research. | RADseq, Genomics, Bioinformatics, Population Genetics | Bioinformatics, Population Genetics | Command Line Tool | Documentation, Uses, and more |
| iqtree | mcc | 1.0 | IQ-TREE is a fast and widely used program for maximum-likelihood phylogenetic inference from molecular sequence alignments (DNA, protein, codon, morphology, and more). It features ModelFinder for automatic selection of the best-fitting substitution model (including partitioned and mixture models), efficient stochastic hill-climbing tree search, the ultrafast bootstrap approximation (UFBoot) and SH-aLRT branch-support tests, and support for partitioned analyses of multi-gene/multi-locus datasets. It also offers tree topology tests, ancestral state reconstruction, and date-calibration features. Input is a sequence alignment (PHYLIP, FASTA, NEXUS) with an optional partition file; outputs include the best ML tree with support values, model-selection reports, and log files. A representative analysis runs ModelFinder Plus followed by many ultrafast bootstrap replicates. This is version 3.0.1 from bioconda. | Phylogenetics, Evolutionary Biology, Bioinformatics | Biochemistry and Molecular Biology, Phylogenetics | Standalone | Documentation, Uses, and more |
| iqtree-2.2.0_beta | mcc | 1.0 | IQ-TREE (version 2.2.0_beta) is an efficient, feature-rich package for maximum-likelihood phylogenetic inference from nucleotide, amino-acid, codon, morphological, or binary alignments. It includes ModelFinder for automatic best-fit substitution-model selection (including partition models and model merging), a fast stochastic tree-search algorithm, the ultrafast bootstrap (UFBoot) and SH-aLRT tests for branch support, and support for partitioned analyses of concatenated multi-locus datasets. Version 2 adds capabilities such as more complex mixture and heterotachy models, tree topology tests (AU, KH, SH), and improved checkpointing. Input is a sequence alignment (PHYLIP/FASTA/NEXUS) plus an optional partition file; outputs include the ML tree in Newick with support values, a report file, and model/parameter estimates. IQ-TREE is one of the most popular tools for building well-supported phylogenies in molecular evolution and systematics. | phylogenetics, bioinformatics, maximum likelihood, tree estimation | Computational Biology, Phylogenetics | Executable | Documentation, Uses, and more |
| isoseq | mcc | 1.0 | Iso-Seq (the PacBio isoseq toolkit) processes PacBio long-read data to characterize full-length transcript isoforms, capturing complete transcripts end to end and thereby resolving alternative splicing, alternative transcription start/end sites, and isoform-level diversity that short reads cannot. The pipeline typically starts from PacBio HiFi/CCS reads: lima (bundled here at 2.9.0) removes primers and demultiplexes, the refine stage trims poly(A) tails and removes concatemers to yield full-length non-chimeric (FLNC) reads, and the cluster/cluster2 stage collapses FLNC reads into high-quality polished transcript consensus sequences; pbskera (1.2.0) is included for deconcatenating segmented (S-reads/MAS-seq) libraries. Outputs are BAM/FASTA files of consensus transcripts that are then mapped to a genome and collapsed into a non-redundant isoform annotation. Version 4.0.0 of isoseq is provided as a conda environment in a Rocky 8 baseline conda container alongside metagenomics/genomics tools such as Kaiju, CellBender, Funannotate, Verkko, BBTools, and seqtk. | RNA-Seq, Transcriptomics, Bioinformatics | Bioinformatics, Molecular Biology | Command Line Tool | Documentation, Uses, and more |
| isoseq3 | mcc, lcc | 1.0 | IsoSeq3 is PacBio's official pipeline for analyzing Iso-Seq (isoform sequencing) data to characterize full-length transcript isoforms without a reference-based assembly. Operating on PacBio HiFi/CCS reads, it removes cDNA primers and demultiplexes barcodes with lima, then its refine step trims poly-A tails and discards chimeras and its cluster step groups full-length non-concatemer reads into polished consensus transcripts, producing high-quality (HQ) and low-quality (LQ) isoform sequences. The result is a set of accurate, full-length transcript models suitable for downstream mapping (e.g. with minimap2) and isoform-level annotation, alternative-splicing, and gene-model refinement studies. In this container it is bundled with its companion PacBio tools lima and pbccs (CCS/HiFi read generation), and the typical flow runs CCS generation, then lima, then IsoSeq refine, then IsoSeq cluster. It is the standard route for turning PacBio long-read cDNA data into curated transcript catalogs. | RNA-Seq, Transcriptomics, PacBio, Isoform Analysis | Molecular Biology, Biological Sciences | Bioinformatics Tool | Documentation, Uses, and more |
| itac | mcc, lcc | N/A | The Intel® Trace Analyzer and Collector profiles and analyzes MPI applications to help focus your optimization efforts. Description Source: https://www.intel.com/content/www/us/en/developer/tools/oneapi/trace-analyzer.html#gs.4oj2fp |
Performance Analysis, Parallel Applications, Optimization | Computer & Information Sciences | Compiler | Documentation, Uses, and more |
| jags | lcc | 1.0 | JAGS (Just Another Gibbs Sampler, conda-forge 4.3.0) is a program for Bayesian hierarchical statistical modeling via Markov Chain Monte Carlo (MCMC), using a dialect of the BUGS model-description language. The user writes a declarative model file specifying stochastic and deterministic relationships among parameters and data (priors, likelihoods, transformations), and JAGS automatically constructs an appropriate Gibbs/Metropolis sampler to draw from the joint posterior distribution. It is engine-agnostic to platform and is most often driven from R through interface packages such as rjags, R2jags, or runjags, which pass data and monitor parameters and return posterior samples for convergence diagnostics (trace plots, Gelman-Rubin) and inference. Inputs are the BUGS-language model, observed data, and initial values; outputs are MCMC samples of monitored nodes. It is widely used across ecology, epidemiology, psychometrics, and other fields for fitting custom hierarchical/mixed models where closed-form inference is intractable. | Bayesian Statistics, Hierarchical Modeling, Mcmc Simulation | Statistics & Probability, Mathematics | Statistical Modeling | Documentation, Uses, and more |
| jasper | mcc | N/A | Jasper is an open-source JPEG image codec library that provides encoding and decoding functionality for JPEG images. It is designed for high performance and is commonly used in applications that require image compression and decompression capabilities. | Image Processing, Open-Source, Jpeg-2000 | Computer Science, Computer & Information Sciences | Image Processing | Documentation, Uses, and more |
| java | mcc, lcc | N/A | Java jdk-11.0.2 | Programming Language, Computing Platform | Software Engineering, Computer & Information Sciences | Compiler | Documentation, Uses, and more |
| jax | lcc | 1.0 | JAX is a Python library for high-performance numerical computing and machine-learning research that pairs a NumPy-compatible array API with composable function transformations: grad for automatic differentiation (forward and reverse mode, including higher-order derivatives), jit for just-in-time compilation to fused, optimized code via XLA, vmap for automatic vectorization, and pmap/sharding for parallelism across devices. The same code runs on CPU, GPU, or TPU, and JAX's functional, immutable-array style underpins many modern ML libraries. This deployment is NVIDIA's NGC GPU build (nvcr.io/nvidia/jax:25.01-py3), providing JAX with a jaxlib compiled against CUDA 12.6, cuDNN 9, and NCCL for multi-GPU communication, plus the Flax neural-network library, with core versions pinned so the GPU stack is consistent and ready to use. Typical use launches Python inside the container and runs JAX/Flax training or scientific-computing code that is JIT-compiled and differentiated on NVIDIA GPUs. | Machine Learning, Numerical Computing, Scientific Computing | Artificial Intelligence, Computer & Information Sciences | Machine Learning | Documentation, Uses, and more |
| jellyfish | mcc, lcc | 1.0 | Jellyfish is a fast, memory-efficient tool for counting k-mers (fixed-length substrings) in DNA sequences, using a lock-free parallel hash table to handle very large datasets on multi-core machines. Counting k-mers is a foundational step for many genomics analyses, including genome-size and heterozygosity estimation (feeding GenomeScope), error correction, sequence assembly, and repeat analysis. Its workflow first counts canonical k-mers from reads into a binary count database, then derives a k-mer frequency histogram from that database or dumps the k-mers and their counts. It reads FASTA/FASTQ input and produces a compact binary count database plus queryable histograms and dumps. This is bioconda jellyfish version 2.2.10. | K-Mer Counting, DNA Sequences, Sequence Analysis | Computational Biology, Biological Sciences | Bioinformatics Tool | Documentation, Uses, and more |
| jmodeltest | lcc | N/A | jmodeltest. | phylogenetics, model selection, bioinformatics, evolutionary biology | Phylogenetics, Evolutionary Biology | Standalone | Documentation, Uses, and more |
| json-c | mcc | N/A | JSON-C implements a reference counting object model that allows you to easily construct JSON objects in C, output them as JSON formatted strings and parse JSON formatted strings back into the C representation of JSON objects. It aims to conform to RFC 7159. Description Source: https://github.com/json-c/json-c/wiki |
C Library, Json Encoding, Json Decoding, Data Manipulation | Computer Science, Computer & Information Sciences | Data Processing | Documentation, Uses, and more |
| jsoncpp | lcc | N/A | JsonCpp is a C++ library that allows manipulating JSON values, including serialization and deserialization to and from strings. It can also preserve existing comment in unserialization/serialization steps, making it a convenient format to store user input files. Description Source: https://github.com/open-source-parsers/jsoncpp |
C++, Json, Library, Data Manipulation | Software Engineering, Computer & Information Sciences | Development Tools | Documentation, Uses, and more |
| juicebox_env | mcc | 1.0 | This environment bundles the tooling needed to scaffold and manually curate genome assemblies using Hi-C chromatin-contact data. It combines Phase Genomics' juicebox_scripts (utilities that convert assembly and contact data to and from the .assembly/.hic formats Juicebox expects, such as generating a review .assembly from an AGP or FASTA) with the Aiden Lab 3D-DNA pipeline, which uses Hi-C read pairs to orient, order, and join contigs into chromosome-length scaffolds and to detect and correct misjoins. Supporting dependencies matlock (BAM/Hi-C manipulation) and htslib (BAM/CRAM I/O) are included. A typical workflow aligns Hi-C reads to a draft assembly, runs 3D-DNA to produce a candidate scaffolding plus a .hic contact map, loads that map into the Juicebox Assembly Tools GUI for interactive review of the contact matrix, then uses the juicebox_scripts to export the reviewed .assembly back into a corrected FASTA. It is packaged as a conda-based Singularity app within a Rocky 9 genome-assembly/Hi-C container. | Documentation, Uses, and more | |||
| julia | mcc, lcc | N/A | Julia is a high-level, high-performance dynamic programming language for technical computing. | Programming Language, Technical Computing, High-Performance Computing, Scientific Computing, Numerical Analysis | Computer Science, Computer & Information Sciences | Compiler | Documentation, Uses, and more |
| jupyter-ai | mcc, lcc, ecc | 1.0 | Jupyter AI (version 3.0.0, installed with the jupyternaut and magics extras) brings generative-AI capabilities directly into JupyterLab and the classic notebook. It provides two main interfaces: IPython magic commands (%%ai and %ai) that let you send a cell's contents to a large language model and render the response inline as markdown, code, HTML, or an image; and Jupyternaut, a conversational chat assistant docked in the JupyterLab sidebar that can answer questions, generate and explain code, and (with learned context) reference your notebooks and files. Version 3 routes requests through a LiteLLM-based model layer, giving provider-agnostic access to many backends (Anthropic Claude, OpenAI, and numerous others) selectable at runtime with your own API keys. In this deployment it is packaged in an Open OnDemand-safe JupyterLab 4 image with real-time collaboration disabled for reverse-proxy compatibility. Typical use is loading the magics extension and then prompting a chosen model from a notebook cell. | Documentation, Uses, and more | |||
| jupyter-custom-container | mcc | 1.0 | Jupyter-Custom-Container is not a software package but a catalog stub that documents the MCC 'bring-your-own-container' Jupyter interactive app served through the MCC Open OnDemand (OOD) portal at mcc-ood.ccs.uky.edu. It lets a researcher launch a JupyterLab/Notebook session on an MCC compute node where the Python/Jupyter runtime and all scientific libraries come from a user-supplied Singularity/Apptainer container image, rather than a fixed site-provided software stack, giving full control over the software environment while OOD handles SLURM scheduling, resource allocation, and secure browser proxying. Through the OOD web form the user selects the container image path along with requested cores, memory, GPUs, partition, and walltime, then connects to the running notebook from the browser. This registration exists purely so the app is discoverable in the SDS catalog; there is no container image built for it here. It is the appropriate choice when a project needs a reproducible, custom Python environment for interactive analysis on the cluster. | Documentation, Uses, and more | |||
| jupyter-notebook | mcc, lcc | 1.0 | This Jupyter Notebook entry is an Open OnDemand interactive application (jupyter-ood-mcc, served from mcc-ood.ccs.uky.edu) that lets researchers launch a JupyterLab/Notebook session running on MCC compute nodes directly from a web browser, with no manual SSH or SLURM scripting. Through the OOD form the user requests resources (cores, memory, walltime, partition, optionally a GPU) and the portal submits and schedules the interactive job, then provides a one-click link into the live notebook server. It supports the standard Jupyter ecosystem for Python (and other kernels) so users can run interactive data analysis, visualization, and machine-learning workloads with full access to the cluster filesystem and installed environments. In the SDS catalog it is a catalog-only stub (not a built container image) registered purely for discoverability, pointing users to the browser-based app rather than a Singularity image. It is the primary interactive computing entry point for many MCC users. | Documentation, Uses, and more | |||
| jupyterlab | ecc | 1.0 | JupyterLab is a browser-based interactive development environment for notebooks, code, and data, served from this container together with a bundled scientific Python 3.10 stack. The environment includes numpy, pandas, matplotlib, seaborn, scikit-learn and scipy for analysis, PyTorch, TensorFlow and Keras for deep learning, Hugging Face transformers and datasets for NLP, nltk and spacy for text processing, and document tools such as PyMuPDF, pdfminer and pytesseract for extracting text from PDFs and images. It is designed to pair with the Ollama server inside the same container, so notebooks can drive local large language models for experimentation, document question answering, and retrieval-augmented workflows without any external service. Users work in notebooks in the browser; results, plots, and files are saved to their ECC storage. | Interactive Computing, Data Visualization, Jupyter Notebooks, Web-Based Interface, Code Editing | Data Analysis, Computer & Information Sciences | Data Science Tool | Documentation, Uses, and more |
| kaiju | mcc | 1.0 | Kaiju (version 1.10.1 from bioconda) performs fast taxonomic classification of metagenomic sequencing reads by translating them in all six frames and searching against a protein reference database using a modified backward-search over an FM-index (BWT). Because it operates at the protein level, Kaiju can be more sensitive than nucleotide k-mer classifiers for divergent or poorly represented organisms and is well suited to classifying reads against databases derived from NCBI nr, RefSeq, or progenomes. Users first build or download an indexed database and then run the classifier, choosing between MEM (exact) and Greedy (mismatch-tolerant) modes; helper scripts convert output to Krona plots or summary tables at any taxonomic rank. It is commonly paired with nucleotide classifiers like Kraken to cross-validate community composition. Outputs integrate with standard downstream taxonomy-summary and visualization tools. | Bioinformatics, Computational Biology, Genomics, Sequencing Data | Biological Sciences | Sequence Analysis Tool | Documentation, Uses, and more |
| kallisto | mcc, lcc | 1.0 | kallisto performs extremely fast quantification of transcript abundances from bulk RNA-seq reads using pseudoalignment, which determines the compatibility of reads with transcripts (via a colored de Bruijn graph / transcriptome index) without producing full base-level alignments. This makes it orders of magnitude faster than alignment-based quantifiers while retaining accuracy, and it produces estimated counts and TPM per transcript along with bootstrap replicates that quantify inferential uncertainty for downstream differential-expression tools such as sleuth. The workflow is two steps: an index step builds an index from a transcriptome FASTA, then a quant step takes FASTQ reads (paired or single-end) and the index and outputs abundance estimates (abundance.tsv/.h5) and run info. It also supports pseudobam output and, via bustools, single-cell workflows. This is version 0.48.0 from bioconda. | RNA-Seq, Transcriptomics, Bioinformatics, Gene Expression Analysis, Computational Biology | Bioinformatics, Biological Sciences | Bioinformatics Tool | Documentation, Uses, and more |
| kat | mcc, lcc | 1.0 | KAT, the K-mer Analysis Toolkit, is a suite for examining k-mer frequency spectra from sequencing reads and assemblies, used chiefly in genome-assembly QC and pre-assembly characterization. It can estimate genome size, heterozygosity, ploidy, sequencing coverage, error content, and assembly completeness, and it flags contamination and repetitive content by comparing spectra. Core subcommands include hist (k-mer frequency histogram), gcp (GC vs coverage), comp (compare two or more datasets, e.g. reads vs assembly, producing spectra-cn plots), and sect (per-sequence coverage). It reads FASTA/FASTQ (optionally gzipped) and emits matrices plus publication-ready density/spectra plots; a common check compares reads against an assembly to reveal missing or duplicated content. This is version 2.4.2. | bioinformatics, genomics, sequence analysis | Bioinformatics, Genomics | Analysis Tool | Documentation, Uses, and more |
| kbproto | mcc, lcc | N/A | The kbproto package provides the core XKB protocol and extension definitions for usage with X11 protocol libraries. | X11 Protocol, Keyboard Properties | Computer Science, Computer & Information Sciences | Protocol | Documentation, Uses, and more |
| kmc | mcc | 1.0 | KMC is a high-performance, disk-based k-mer counter designed to tally the frequencies of fixed-length substrings (k-mers) in very large DNA sequencing datasets while keeping memory usage bounded by streaming data to and from temporary disk storage. It handles FASTQ/FASTA input (gzipped or not, single or multiple files) and produces a compact KMC database that downstream tools can query, and it is used for genome-size and coverage estimation, error correction, read filtering, and as a preprocessing step for assemblers and metagenomic classifiers. The suite includes kmc for counting and kmc_tools for set operations on databases (intersection, union, subtraction, complex filtering) and for dumping counts to text, and its transform capability can build a k-mer frequency histogram from a counted database. Version 3.2.4 is provided as a conda app in a Rocky 9 genome-assembly/Hi-C container. | Sequence Analysis, Genome Assembly, Metagenomics | Computational Biology, Biological Sciences | Tool | Documentation, Uses, and more |
| kneaddata | mcc, lcc | 1.0 | KneadData is a quality-control pipeline from the Huttenhower lab tailored to metagenomic and metatranscriptomic sequencing, where removing contaminant host and other unwanted reads is essential before taxonomic or functional profiling. It wraps Trimmomatic for adapter and quality trimming and Bowtie2 (or BMTagger) to align reads against one or more contaminant reference databases—commonly the human genome, human transcriptome, or rRNA sequences—discarding reads that map so that only microbial reads remain. It also handles tandem-repeat filtering and produces detailed logs of how many reads were removed at each step. It takes input FASTQ reads and one or more contaminant reference databases, yielding cleaned paired FASTQ files ready for tools like MetaPhlAn or HUMAnN. Version 0.12.2 is the standard preprocessing front-end in the bioBakery microbiome analysis suite. | Bioinformatics, Metagenomics, Data Quality Control | Biomolecular Sciences, Biological Sciences | Analysis Tool | Documentation, Uses, and more |
| kount | lcc | 1.0 | KOunt is a Snakemake-based pipeline for quantifying KEGG Orthologs (KOs) from metagenomic and metatranscriptomic sequencing data, giving a functional-abundance profile of a microbial community. It orchestrates the steps of read QC, assembly, gene prediction, and annotation against KEGG to assign genes to KO identifiers, then counts reads mapped to those genes to produce per-KO abundance tables aggregated across samples, supporting comparative functional metagenomics. As a Snakemake workflow it is configured through a config file specifying inputs and parameters and executed with Snakemake using its conda/mamba-managed environments for reproducibility. It is obtained from the WatsonLab/KOunt GitHub repository. | Documentation, Uses, and more | |||
| kraken-biom | lcc | 1.0 | kraken-biom is a utility that converts the taxonomic-classification reports produced by Kraken (and Kraken2/Bracken and related classifiers) into the BIOM (Biological Observation Matrix) format used by microbiome-analysis ecosystems such as QIIME 2 and phyloseq. It reads one or more Kraken report files — one per sample — and aggregates their per-taxon read counts into a single sample-by-taxon OTU/feature table with an associated taxonomy, so that shotgun-metagenomics classification results can be analyzed with the same downstream diversity, ordination, and differential-abundance tools developed for amplicon data. Inputs are the Kraken report files (plus optional sample metadata); the output is a BIOM table (JSON or HDF5). This is version 1.0.1 from bioconda. | Documentation, Uses, and more | |||
| kraken2 | mcc | 1.0 | Kraken 2 (version 2.1.2 from bioconda) is a fast taxonomic classifier that assigns taxonomic labels to DNA sequences by matching their constituent k-mers against a database mapping k-mers to the lowest common ancestor of the genomes that contain them, using a compact hash for high speed and reduced memory versus the original Kraken. Users build or download a database (bacterial/archaeal/viral RefSeq, standard, or custom) with the kraken2-build utility, then classify reads against that database to obtain per-read assignments and a hierarchical report of read counts per taxon. Its confidence scoring and report output pair naturally with Bracken for species-level abundance re-estimation and with Krona for visualization. It is a cornerstone of metagenomic community profiling and contamination screening, valued for classifying millions of reads in minutes. Outputs feed directly into downstream abundance and diversity analyses. | bioinformatics, sequence classification, metagenomics, taxonomy | Bioinformatics, Metagenomics | Command-line tool | Documentation, Uses, and more |
| kraken2-2.1.3 | mcc | 1.0 | Kraken2 (bioconda 2.1.3) is a fast, exact-k-mer-based taxonomic classification system that assigns taxonomic labels to DNA sequencing reads by matching their constituent k-mers against a database mapping k-mers to the lowest common ancestor (LCA) of the genomes containing them. Version 2 uses a compact, probabilistic hashed representation of minimizers, dramatically reducing memory and disk requirements relative to the original Kraken while maintaining high classification speed and accuracy. It supports building custom databases (via kraken2-build, drawing from RefSeq libraries or user genomes) and classifying single- or paired-end reads, producing per-read assignments and a summary report of read counts per taxon. Its reports are commonly refined with Bracken for abundance re-estimation and visualized with tools like Krona or Pavian; it is a workhorse for metagenomic profiling and host/contaminant screening. | bioinformatics, sequence classification, metagenomics | Bioinformatics, Biological Sciences | Command-line tool | Documentation, Uses, and more |
| krakenuniq | mcc | 1.0 | KrakenUniq is a metagenomic taxonomic classifier that extends the exact k-mer matching approach of Kraken with an additional count of the number of unique k-mers observed per taxon, computed efficiently using the HyperLogLog cardinality-estimation algorithm. This unique-k-mer count sharply reduces false-positive species identifications common in low-abundance or contaminated samples, because a genuinely present organism should contribute many distinct k-mers rather than the same few repeatedly. It reads sequencing reads (FASTA/FASTQ) against a prebuilt k-mer-to-taxon database and outputs per-read classifications plus a report giving, for each taxon, both read counts and estimated unique-k-mer coverage. It is backwards-compatible with Kraken 1 databases while adding the coverage-aware report. This is bioconda krakenuniq 1.0.4. | bioinformatics, metagenomics, sequence classification | Bioinformatics, Biological Sciences | Command-line tool | Documentation, Uses, and more |
| krb5 | mcc | N/A | Kerberos is a network authentication protocol designed to provide strong authentication for client/server applications through secret-key cryptography. | Authentication, Security, Networking | Computer Science | Library | Documentation, Uses, and more |
| krona | mcc | 1.0 | Krona (version 2.8.1) creates interactive, zoomable HTML visualizations of hierarchical data, most commonly the taxonomic composition of metagenomic samples. Its signature output is a radial, space-filling (multi-level pie/sunburst) chart in which each ring is a taxonomic rank; users can click any wedge to zoom into a clade, adjust the depth displayed, and hover for counts and percentages, all in a self-contained HTML file that opens in any web browser without a server. Krona includes importers that convert output from classifiers and BLAST-style results into charts, such as an import-taxonomy tool, an import-BLAST tool, and an import-text tool for generic tab-delimited hierarchies. Because it makes complex community-composition data easy to explore and share, Krona is a near-ubiquitous reporting step in metagenomics and amplicon pipelines (e.g. downstream of Kraken, Centrifuge, or QIIME). | Documentation, Uses, and more | |||
| lammps | mcc, lcc | 1.0 | Large-scale Atomic/Molecular Massively Parallel Simulator | Molecular Dynamics, Simulation, Computational Chemistry | Physical Chemistry, Chemical Sciences | Code/Package | Documentation, Uses, and more |
| lammps-kokkos-cufft-h200 | lcc | 1.0 | LAMMPS (Large-scale Atomic/Molecular Massively Parallel Simulator, built here from the lammps GitHub repo, branch stable_22Jul2025_update4) is a classical molecular-dynamics engine for simulating systems from atoms and coarse-grained beads to mesoscale particles across materials science, chemistry, and soft matter. This build targets GPU acceleration through the KOKKOS CUDA backend and uses cuFFT to perform the long-range electrostatics FFTs (KSPACE/PPPM) on the GPU, is compiled for NVIDIA H200 (Hopper, sm_90), and is MPI-enabled for multi-GPU/multi-node scaling on LCC. It includes a broad set of physics packages such as MOLECULE, KSPACE, MANYBODY, REAXFF (reactive force fields), and ML-SNAP (machine-learned potentials), enabling everything from biomolecular and polymer simulations to reactive and machine-learning-potential runs. Users drive it with an input script, launching under MPI with the KOKKOS GPU backend enabled. Trajectories and thermodynamic output are written for downstream analysis in tools like OVITO or VMD. | Documentation, Uses, and more | |||
| lammps-kokkos-v100 | lcc | 1.0 | LAMMPS (Large-scale Atomic/Molecular Massively Parallel Simulator) is a classical molecular-dynamics code from Sandia National Laboratories for modeling materials at the atomic and mesoscopic scale, from biomolecules and polymers to metals, semiconductors, and coarse-grained systems. This build is the stable_22Jul2025_update3 release compiled with the KOKKOS acceleration package for CUDA, specifically targeting NVIDIA V100 (Volta, sm_70) GPUs, and is MPI-enabled for multi-node scaling. It bundles a broad set of physics packages including MOLECULE, KSPACE (long-range electrostatics), MANYBODY (Tersoff/Stillinger-Weber style potentials), REAXFF (reactive force fields), and ML-SNAP (machine-learned potentials), covering a wide range of interatomic models. Simulations are driven by an input script that defines the system, force field, and run commands, with output in dump/thermo files for trajectory and property analysis. The KOKKOS-CUDA backend offloads force computation to the GPU for large speedups on Volta hardware. | Documentation, Uses, and more | |||
| lammps-mace | lcc | 1.0 | This is a GPU-accelerated build of LAMMPS (Large-scale Atomic/Molecular Massively Parallel Simulator), the classical molecular-dynamics engine widely used across materials science, chemistry, and soft-matter physics, from the ACEsuit mace fork that wires in the MACE machine-learning interatomic potential. MACE is a higher-order equivariant message-passing neural-network potential that delivers near-DFT accuracy at a fraction of the cost, evaluated inside LAMMPS through the pair_style mace interface backed by LibTorch. This particular build uses the Kokkos CUDA 12.4 acceleration package compiled for the H200 (compute capability sm_90) on LCC, so it runs the potential on-GPU. A user supplies a standard LAMMPS input script that loads a trained MACE model file together with a data/geometry file, selecting the MACE pair style and pointing it at the model, and launches with the appropriate Kokkos GPU acceleration flags. It supports the usual LAMMPS MD workflows — NVE/NVT/NPT ensembles, minimization, and trajectory/thermo output — but with an ML potential in place of a classical force field. | Documentation, Uses, and more | |||
| lammps-mace-mliap | lcc | 1.0 | This is a GPU-accelerated build of LAMMPS coupled to MACE machine-learning interatomic potentials for atomistic and molecular-dynamics simulation. LAMMPS (Large-scale Atomic/Molecular Massively Parallel Simulator) is a classical MD engine that integrates Newton's equations for systems from a few atoms to billions, and here it is compiled from the develop branch with the ML-IAP Python interface plus the Kokkos/CUDA acceleration package targeting modern NVIDIA GPUs (H200-class). MACE is an equivariant message-passing graph neural network that predicts DFT-quality energies and forces; installed via mace-torch (with CUDA 12.4 PyTorch), its trained models are used inside LAMMPS as a machine-learning pair style, letting users run large near-first-principles MD without hand-tuned force fields. The workflow is standard LAMMPS: an input script defines the system and simulation, and a pair style referencing a MACE model file drives the energetics; foundation models or user-trained potentials can both be supplied. Built from LAMMPS develop plus mace-torch 0.3.15. | Documentation, Uses, and more | |||
| lammpss-deepmd-kit-module-cpu | lcc | 1.0 | This is a CPU build of LAMMPS (Large-scale Atomic/Molecular Massively Parallel Simulator) coupled with the DeePMD-kit plugin, enabling classical molecular-dynamics simulations that use machine-learned 'deep potential' interatomic models in place of, or alongside, conventional analytic force fields. DeePMD-kit provides neural-network potentials trained on quantum-mechanical (DFT) reference data that reproduce ab-initio accuracy at far lower cost, and the integration exposes them to LAMMPS through a deepmd pair style so that a trained model (a frozen graph) drives the forces during dynamics. Beyond running MD, DeePMD-kit supports training, freezing, and compressing models and testing them, while LAMMPS handles the ensemble integration, thermostats/barostats, and standard MD analysis. Simulations are configured with an ordinary LAMMPS input script that specifies the deepmd pair style and the model file, and are run (with MPI for parallelism) via the LAMMPS executable. Because this is the CPU-only conda build (from the deepmodeling channel), it runs without a GPU, trading throughput for portability; it is used for materials and molecular simulations requiring near-DFT accuracy. | Documentation, Uses, and more | |||
| lammpss-deepmd-kit-module-gpu | lcc | 1.0 | This deployment couples LAMMPS with the DeePMD-kit plugin in a GPU-enabled build, enabling molecular-dynamics simulations that use machine-learned deep-potential interatomic models in place of hand-crafted classical force fields. DeePMD-kit trains deep neural-network potentials on reference quantum-mechanical (DFT) energies and forces, achieving near-DFT accuracy at a tiny fraction of the cost, and the LAMMPS integration (lammps-dp) exposes those trained models through a deep-potential pair style so that standard LAMMPS input scripts can run dynamics driven by the ML potential on the GPU. The workflow is typically to prepare training data, train and freeze a model with the DeePMD-kit tooling, then run production MD in LAMMPS calling that frozen model. It is used in computational materials science and chemistry to simulate systems (reactive, phase-change, catalytic) where classical potentials are inadequate but full ab initio MD would be prohibitively expensive. | Documentation, Uses, and more | |||
| laser | mcc | 1.0 | LASER (Locating Ancestry from SEquence Reads) is a method for estimating an individual's genetic ancestry by projecting their sequence data into the principal-component (PCA) coordinate space defined by a reference panel of genotyped individuals. Its distinctive feature is that it works directly from sequence reads — including low-coverage or off-target/shotgun data such as the incidental human reads in a sequencing experiment — by simulating the same read sampling process for reference individuals so that the study sample and references are placed on a common PCA map despite differing data quality. This enables ancestry inference and identification of population structure without requiring high-quality genotype calls. The workflow takes a pileup/sequence-based representation of the sample together with a reference genotype panel (e.g. HGDP or 1000 Genomes coordinates) and outputs the sample's projected PCA coordinates, which are then interpreted against the reference populations. The related TRACE program in the package handles genotype (rather than read) input. This deployment installs LASER version 2.04 from a locally staged tarball. | Documentation, Uses, and more | |||
| lastz | lcc | 1.0 | LASTZ (version 1.0.4) is a pairwise DNA sequence aligner developed as a successor to BLASTZ, designed for comparing large genomic sequences and whole chromosomes. It performs sensitive seed-and-extend local alignment with configurable seeding patterns, scoring matrices, and gapped-extension parameters, and is well suited to aligning divergent or repeat-rich genomic regions. It is a common first step in whole-genome comparative and synteny analyses — for instance generating the alignments that are chained and netted to build UCSC-style chain/net files and liftOver chains. Inputs are two FASTA sequences (a target and a query, with options to mask or soft-mask repeats), and output can be produced in several formats including general tab-delimited hits, MAF, SAM, and dot-plot data for visualization, with numerous flags to tune sensitivity and speed. It is widely used in comparative genomics and genome-assembly comparison. | Bioinformatics, DNA Sequencing, Alignment, Comparative Genomics | Genomics, Biological Sciences | Bioinformatics Tool | Documentation, Uses, and more |
| latex | lcc | N/A | LaTeX (provided via TeX Live 2020) is a document preparation and typesetting system built on TeX, the standard for producing high-quality scientific and technical documents. It excels at typesetting mathematics, bibliographies, cross-references, and structured documents such as papers, theses, and books. Researchers across all disciplines use it to author publications and reports. | Typesetting, Scientific Communication, Document Preparation | Typesetting and Formatting, Natural Sciences, Mathematics, Computer & Information Sciences, Engineering & Technology | Publishing | Documentation, Uses, and more |
| lcc-user | lcc | N/A | Documentation, Uses, and more | ||||
| lcc-users | lcc | N/A | Documentation, Uses, and more | ||||
| lcms | mcc | N/A | LCMS (Liquid Chromatography Mass Spectrometry) is a software package designed for the analysis of mass spectrometry data, particularly in the context of liquid chromatography. | Mass Spectrometry, Data Analysis, Chemistry, Bioinformatics | Chemistry, Chemical Sciences | Data Analysis | Documentation, Uses, and more |
| lefse | mcc | 1.0 | LEfSe (Linear discriminant analysis Effect Size) is a biomarker-discovery method widely used in microbiome and metagenomics studies to find features (taxa, genes, pathways, or other measured quantities) that both differ significantly between two or more biological classes and carry substantial effect size. It couples a non-parametric Kruskal-Wallis rank-sum test across classes with pairwise Wilcoxon tests among subclasses, and then estimates the magnitude of each significant difference using Linear Discriminant Analysis, reporting an LDA effect-size score per feature. Input is a tab-delimited abundance table with class (and optional subclass/subject) rows; the pipeline first formats and prepares the data, then performs the statistics, and finally uses plotting utilities to render ranked bar charts and taxonomic cladograms highlighting the discriminative features. This is the Galaxy/bioconda packaging (v1.1.2) provided as a conda app in a large multi-tool genomics container on Rocky 8. | bioinformatics, microbiome, statistical analysis, biomarker discovery | Computational Biology, Biological Sciences | Analysis Tool | Documentation, Uses, and more |
| lftp | mcc | 1.0 | lftp is a sophisticated command-line file-transfer client that supports many protocols — FTP, FTPS, HTTP, HTTPS, SFTP, FISH, and more — behind a single interactive shell-like interface, and is favored for reliable bulk and automated transfers. Its standout features include robust automatic retry and resume of interrupted transfers, background/queued jobs, and segmented (multi-connection) downloads via its pget command that can dramatically speed up large files. Its mirror command recursively synchronizes whole directory trees in either direction (with a reverse mode to upload, parallel transfers for concurrency, and options to delete extraneous files or only transfer newer ones), making it a common choice for pushing/pulling large datasets to and from remote servers and data repositories. lftp can be driven interactively or scripted non-interactively by passing commands or a script file, and it handles authentication, bookmarks, and rate limiting. This build is version 4.9.2 from conda-forge. | File Transfer | Networking, Computer Science | Command Line Tool | Documentation, Uses, and more |
| lhapdf | lcc | 1.0 | LHAPDF is the standard C++ library (with Python bindings) in high-energy particle physics for evaluating parton distribution functions (PDFs) and their associated alpha_s running-coupling values, providing uniform programmatic access to the many PDF set grids used in QCD and collider phenomenology. It loads interpolation grids from installed PDF sets and returns x*f(x, Q^2) values for given parton flavors, momentum fraction x, and energy scale Q, supporting error-set uncertainty evaluation across a set's members. It is used from C++ or Python within Monte Carlo generators and analysis code. This is version 6.5.5, provided with JupyterLab and matplotlib so PDFs can be evaluated and plotted interactively; note that specific PDF set data files must be installed separately. | High-Energy Physics, Parton Distribution Functions, Computational Physics | High-Energy Physics, Particle and High-Energy Physics | Library | Documentation, Uses, and more |
| libbsd | mcc, lcc, ecc | N/A | libbsd is a library that provides a collection of functions from the FreeBSD operating system for enhancing compatibility and portability across different Unix-like systems. It includes various system-related functions that are widely used in Unix programming to access system services and perform low-level operations. | Library, Bsd, Operating System | Software Engineering, Computer & Information Sciences | Library | Documentation, Uses, and more |
| libcroco | mcc | N/A | libcroco is a library for manipulating CSS (Cascading Style Sheets) in a flexible and efficient manner, primarily used in web development and rendering engines. | CSS, Web Development, Library, Rendering | Software Engineering, Other Computer and Information Sciences | Library | Documentation, Uses, and more |
| libdrm | mcc | N/A | This is libdrm, a userspace library for accessing the DRM, direct rendering manager, on Linux, BSD and other operating systems that support the ioctl interface. The library provides wrapper functions for the ioctls to avoid exposing the kernel interface directly, and for chipsets with drm memory manager, support for tracking relocations and buffers. Description Source: https://gitlab.freedesktop.org/mesa/drm |
Graphics, Linux, Kernel, Driver | Software Engineering, Computer Science | Library | Documentation, Uses, and more |
| libedit | mcc, lcc | N/A | libedit is a library for reading and editing user input in a command-line interface. It provides tools for handling input editing, command history, and line parsing in a terminal environment. | Library, Command-Line, Line Editing, History, Open-Source | Software Engineering, Computer & Information Sciences | Tool | Documentation, Uses, and more |
| libepoxy | mcc | N/A | Epoxy is a library for handling OpenGL function pointer management for you. It hides the complexity of dlopen(), dlsym(), glXGetProcAddress(), eglGetProcAddress(), etc. from the app developer, with very little knowledge needed on their part. They get to read GL specs and write code using undecorated function names like glCompileShader(). Description Source: https://github.com/anholt/libepoxy |
Opengl, Library, Graphics | Computer Science, Computer & Information Sciences | Development | Documentation, Uses, and more |
| libevent | mcc, lcc | N/A | Libevent is an event notification library for developing scalable network servers. The Libevent API provides a mechanism to execute a callback function when a specific event occurs on a file descriptor or after a timeout has been reached. Furthermore, Libevent also support callbacks due to signals or regular timeouts. Description Source: https://libevent.org/doc/ |
Event Notification, Callback Functions, Scalable, Efficient | Software Engineering, Computer & Information Sciences | Development | Documentation, Uses, and more |
| libexif | mcc | N/A | libexif is a library for reading and writing EXIF (Exchangeable Image File Format) data in image files. | Camera make, Camera model, Exposure time, F-number, ISO speed ratings, Date and time | Software Engineering, Computer and Information Sciences | Library | Documentation, Uses, and more |
| libfabric | mcc, lcc | N/A | Development files for the libfabric library | Fabric Interfaces, High-Performance Computing, Communication Resources | Computer Science, Engineering & Technology | Communication Library | Documentation, Uses, and more |
| libffi | mcc, lcc, ecc | N/A | The libffi library provides a portable, high level programming interface to various calling conventions. This allows a programmer to call any function specified by a call interface description at run-time. FFI stands for Foreign Function Interface. A foreign function interface is the popular name for the interface that allows code written in one language to call code written in another language. The libffi library really only provides the lowest, machine dependent layer of a fully featured foreign function interface. Description Source: https://github.com/libffi/libffi |
Library, Portability, Calling Conventions | Software Engineering, Computer & Information Sciences | Development Tool | Documentation, Uses, and more |
| libfontenc | mcc | N/A | libfontenc is a library used for encoding font data in X Window System, allowing for the representation of various character sets. | Font Management, X Window System, Character Encoding | Software Engineering, Other Computer and Information Sciences | Library | Documentation, Uses, and more |
| libfuse | mcc, ecc | N/A | libfuse is a library that allows the creation of fully functional file systems in userspace. It provides a simple interface for implementing file systems without the need for kernel-level code. | File System, Userspace, Linux, Development | Software Engineering, Computer Science | Library | Documentation, Uses, and more |
| libgcrypt | mcc | N/A | Libgcrypt is a general purpose cryptographic library originally based on code from GnuPG | Cryptography, Security, Encryption | Computer Science, Computer & Information Sciences | Cryptographic Library | Documentation, Uses, and more |
| libgit2 | mcc | N/A | libgit2 is a portable, pure C implementation of the Git core methods provided as a re-entrant linkable library with a solid API, allowing you to write native speed custom Git applications in any language that supports C bindings. Description Source: https://libgit2.org/ |
Version Control System, Git, Library | Software Engineering, Computer & Information Sciences | Version Control System | Documentation, Uses, and more |
| libgpg-error | mcc, ecc | N/A | libgpg is a library that provides an API for interacting with GnuPG (GNU Privacy Guard) for encryption, decryption, and key management tasks. It allows developers to integrate GnuPG functionality into their applications. | Error Handling, Library, Gnupg | Software Engineering, Computer & Information Sciences | Utility | Documentation, Uses, and more |
| libice | mcc, lcc | N/A | libice is a library that provides support for Inter-Client Exchange (ICE) in the X Window System, allowing X clients to securely exchange data through the X server. | Library, Inter-Client Exchange, X Window System | Computer & Information Sciences | Communication | Documentation, Uses, and more |
| libiconv | mcc, lcc, ecc | N/A | libiconv provides an iconv() implementation, for use on systems which don't have one, or whose implementation cannot convert from/to Unicode. Description Source: https://www.gnu.org/software/libiconv/#TOCintroduction |
Character Set Conversion, Text Data, Encoding | Other Computer & Information Sciences, Computer & Information Sciences | Conversion Tool | Documentation, Uses, and more |
| libid3tag | mcc | N/A | Documentation, Uses, and more | ||||
| libidn2 | mcc | N/A | libidn2 is a library that implements internationalized domain name (IDN) encoding and decoding functions. It provides support for handling domain names with non-ASCII characters according to the IDNA2008 specification. | Domain Names, Idns, Stringprep, Punycode | Software Engineering, Computer & Information Sciences | Utility | Documentation, Uses, and more |
| libint | mcc, lcc | N/A | libint is a library for the evaluation of molecular integrals of many-body operators over Gaussian functions. Description Source: https://github.com/evaleev/libint?tab=readme-ov-file |
Computational Chemistry, Molecular Integrals, Quantum Chemistry | Physical Chemistry, Chemical Sciences | Computational | Documentation, Uses, and more |
| libjpeg-turbo | mcc, lcc | N/A | libjpeg-turbo is a JPEG image codec that uses SIMD instructions (MMX, SSE2, AVX2, Neon, AltiVec) to accelerate baseline JPEG compression and decompression on x86, x86-64, Arm, and PowerPC systems, as well as progressive JPEG compression on x86, x86-64, and Arm systems. On such systems, libjpeg-turbo is generally 2-6x as fast as libjpeg, all else being equal. On other types of systems, libjpeg-turbo can still outperform libjpeg by a significant amount, by virtue of its highly-optimized Huffman coding routines. Description Source: https://libjpeg-turbo.org/ |
Image Compression, Jpeg Compression, Library, Open Source | Image Processing, Computer & Information Sciences | Compression Library | Documentation, Uses, and more |
| libmd | mcc, lcc, ecc | N/A | libmd is a library for molecular dynamics simulations that allows for the integration of equations of motion for particles interacting with different potentials. | Molecular Dynamics, Simulation, Library | Physical Sciences | Computational | Documentation, Uses, and more |
| libnl | mcc | N/A | libnl is a set of libraries providing a generic netlink communication protocol for Linux kernel networking subsystems. It simplifies the process of interacting with the kernel networking stack. | Networking, Linux, Netlink, Kernel | Computer Science | Library | Documentation, Uses, and more |
| libnsl | mcc | N/A | Public client interface library for NIS(YP) Description Source: https://github.com/conda-forge/libnsl-feedstock/blob/main/recipe/meta.yaml | Network Services, Communication, Protocols | Computer Science, Computer & Information Sciences | System Software | Documentation, Uses, and more |
| libogg | lcc | N/A | libogg is a multimedia container format, and the native file and stream format for the Xiph.org multimedia codecs. As with all Xiph.org technology is it an open format free for anyone to use. Description Source: https://www.xiph.org/ogg/ |
Library, Bitstream, Audio | Computer Science, Computer & Information Sciences | Data Format | Documentation, Uses, and more |
| libpciaccess | mcc, lcc, ecc | N/A | libpciaccess is a generic PCI access library. Description Source: https://cgit.freedesktop.org/xorg/lib/libpciaccess |
Library, Pci Bus Access, User-Space Drivers | Software Engineering, Computer & Information Sciences | Programming Library | Documentation, Uses, and more |
| libpng | mcc, lcc | N/A | libpng is the official PNG reference library. It supports almost all PNG features, is extensible, and has been extensively tested for over 28 years. Description Source: http://www.libpng.org/pub/png/libpng.html |
Image Processing, Graphics Library, Raster Images | Image Processing, Computer & Information Sciences | Data Handling | Documentation, Uses, and more |
| libpthread-stubs | mcc, lcc | N/A | The libpthread-stubs package provides weak aliases for pthread functions not provided in libc or otherwise available by default. | Documentation, Uses, and more | |||
| librsvg | mcc | N/A | librsvg is a library to render SVG images to Cairo surfaces. GNOME uses this to render SVG icons. Outside of GNOME, other desktop environments use it for similar purposes. Wikimedia uses it for Wikipedia's SVG diagrams. Description Source: https://gitlab.gnome.org/GNOME/librsvg |
Software Library, Free Software, Svg Rendering, File Conversion | Software Engineering, Computer & Information Sciences | Rendering & Conversion | Documentation, Uses, and more |
| libseccomp | ecc | N/A | libseccomp is a library for implementing secure computing mode (seccomp) filters in Linux systems. It enables developers to apply fine-grained filtering on system calls, restricting the actions that processes can perform for improved system security. | Syscall Filtering, Linux Kernel, Security | Systems Security, Computer & Information Sciences | Development | Documentation, Uses, and more |
| libsigsegv | mcc, lcc, ecc | N/A | libsigsegv is a library for handling segmentation fault signals in Unix-like systems. It provides a mechanism for detecting, handling, and recovering from segment violation errors during program execution. | Library, Memory, Page Fault | Computer Science, Computer & Information Sciences | Utility | Documentation, Uses, and more |
| libsm | mcc, lcc | N/A | Libsm is a library that provides a set of common utilities for scientific computing in C++. | Library, Scientific Computing, C++ | Other Natural Sciences | Library | Documentation, Uses, and more |
| libssh2 | mcc | N/A | libssh2 is a library implementing the SSH2 protocol, providing a way to securely connect to remote servers and execute commands or transfer files. | SSH, Networking, Security, Library | Computer Science | Library | Documentation, Uses, and more |
| libtheora | lcc | N/A | libtheora is a free and open-source software library for encoding and decoding video in the Theora format, which is a lossy video compression format. | video, codec, open-source, multimedia | Computer Science, Computer and Information Sciences | Library | Documentation, Uses, and more |
| libtiff | mcc, lcc | N/A | The LibTIFF software provides support for the Tag Image File Format (TIFF), a widely used format for storing image data. Description Source: http://www.simplesystems.org/libtiff/ |
Image Processing, File Format, Library, Tiff, Image Storage | Computer Science, Computer & Information Sciences | Image Processing | Documentation, Uses, and more |
| libtirpc | mcc | N/A | The libtirpc package contains libraries that support programs that use the Remote Procedure Call (RPC) API. It replaces the RPC, but not the NIS library entries that used to be in glibc. Description Source: https://www.linuxfromscratch.org/blfs/view/stable/basicnet/libtirpc.html |
Software Library | Computer Science, Computer & Information Sciences | Development | Documentation, Uses, and more |
| libtool | mcc, lcc, ecc | N/A | autoamke software. | Library Management, Software Development, Compilation | Software Engineering, Engineering & Technology | Utility/Tool | Documentation, Uses, and more |
| libunistring | mcc | N/A | libunistring is a library that provides functions for manipulating Unicode strings and characters, offering support for various Unicode operations and text processing tasks. It enables developers to work with Unicode data in a portable and efficient manner. | Unicode, Encoding, String Manipulation, Localization | Software Engineering, Computer & Information Sciences | Computational Software | Documentation, Uses, and more |
| libunwind | lcc | N/A | libunwind is a portable and efficient C API for determining the current call chain of ELF program threads of execution and for resuming execution at any point in that call chain. The API supports both local (same process) and remote (other process) operation. Description Source: https://github.com/libunwind/libunwind |
Debugging, Profiling, Runtime Instrumentation, Call Stack, C Programming | Software Engineering, Engineering & Technology | Development | Documentation, Uses, and more |
| libx11 | mcc, lcc | N/A | libx11 is the X11 protocol client library. Description Source: https://gitlab.freedesktop.org/xorg/lib/libx11 |
X Window System, Client Library, Window System Interface, Graphics | Software Engineering, Computer & Information Sciences, Computer Science | Software Development | Documentation, Uses, and more |
| libxau | mcc, lcc | N/A | libXau is a library that provides authentication data for Xlib-based client programs. | Library, Authentication, Xlib | Computer & Information Sciences | Documentation, Uses, and more | |
| libxc | mcc, lcc | N/A | libxc is a library of exchange-correlation and kinetic energy functionals for density-functional theory. The original aim was to provide a portable, well tested and reliable set of these functionals to be used by all the codes of the European Theoretical Spectroscopy Facility (ETSF), but the library has since grown to be used in several other types of codes as well; see below for a partial list. Description Source: https://www.tddft.org/programs/libxc/ |
Quantum Mechanics, Computational Chemistry | Theoretical Chemistry, Chemical Sciences | Density Functional Theory | Documentation, Uses, and more |
| libxcb | mcc, lcc | N/A | The X protocol C-language Binding (libXCB) is a replacement for Xlib featuring a small footprint, latency hiding, direct access to the protocol, improved threading support, and extensibility. Description Source: https://xcb.freedesktop.org/ |
X Window System, C Library, X Protocol | Computer Science, Computer & Information Sciences | Programming | Documentation, Uses, and more |
| libxcrypt | ecc | N/A | libxcrypt is a library for cryptographic hashing and password hashing, providing modern and secure algorithms for password storage and verification. | Cryptography, Security, Password Hashing, Libraries | Cryptography, Computer Science | Library | Documentation, Uses, and more |
| libxdamage | mcc | N/A | Documentation, Uses, and more | ||||
| libxdmcp | mcc, lcc | N/A | libxdmcp is the X Display Manager Control Protocol library. It provides an API for the X Display Manager Control Protocol (XDMCP) for both the client and server side. | X11, Display Manager, Network Protocol | Software Engineering, Computer & Information Sciences | Networking | Documentation, Uses, and more |
| libxext | mcc, lcc | N/A | libXext is an extension library for the X Window System that provides additional functionality to the core Xlib library. | X Window System, Graphics, Library | Computer Science, Computer and Information Sciences | Library | Documentation, Uses, and more |
| libxfixes | mcc | N/A | libxfixes is a library that provides an interface for the X Window System to manage and manipulate window properties and visual effects. | X Window System, Graphics, Window Management | Software Engineering, Computer Science | Library | Documentation, Uses, and more |
| libxfont | mcc | N/A | libXfont is a library used for font management in the X Window System, providing functionalities to load and manage bitmap and scalable fonts. | X Window System, Font Management, Graphics | Computer Science, Software Engineering | Library | Documentation, Uses, and more |
| libxft | mcc | N/A | libXft is a library for rendering text in X Window System applications using the FreeType font rendering library. | X11, Font Rendering, Graphics, Open Source | Software Engineering, Computer Science | Library | Documentation, Uses, and more |
| libxi | mcc | N/A | Documentation, Uses, and more | ||||
| libxinerama | mcc | N/A | libXinerama is a library that provides an interface for the X Window System to support multiple screens and display configurations. | X Window System, Multi-screen, Graphics | Software Engineering, Computer Science | Library | Documentation, Uses, and more |
| libxkbcommon | mcc | N/A | libxkbcommon is a library that provides an interface to the X Keyboard Extension (XKB) for handling keyboard input in a platform-independent manner. It is designed to be used in applications that need to manage keyboard layouts, keymaps, and input events. | keyboard, input handling, X11, cross-platform | Software Engineering, Other Computer and Information Sciences | Library | Documentation, Uses, and more |
| libxkbfile | mcc | N/A | libxkbfile is a library for reading and writing XKB (X Keyboard Extension) files. It provides a simple API for manipulating keyboard configuration files used in X Window System. | X Window System, Keyboard Configuration, Library | Software Engineering, Other Computer and Information Sciences | Library | Documentation, Uses, and more |
| libxml2 | mcc, lcc, ecc | N/A | libxml2 is an XML toolkit implemented in C, originally developed for the GNOME Project. Description Source: https://gitlab.gnome.org/GNOME/libxml2/-/blob/master/README.md?ref_type=heads |
Xml, Html, Parsing, Validation, Manipulation | Software Engineering, Computer & Information Sciences | Parsing | Documentation, Uses, and more |
| libxmu | mcc | N/A | libXmu is a library that provides various utility functions for X Window System applications, particularly those using the X Toolkit. | X Window System, Graphics, Utilities, Libraries | Computer Science, Computer and Information Sciences | Library | Documentation, Uses, and more |
| libxrandr | mcc, lcc | N/A | libxrandr is a library that provides an interface to the X Resize and Rotate (RandR) extension of the X Window System, allowing applications to dynamically change the size, orientation, and reflection of the display. | X11, Graphics, Display Management | Software Engineering, Computer Science | Library | Documentation, Uses, and more |
| libxrender | mcc, lcc | N/A | libXrender is a library that provides an interface to the X Rendering Extension, which allows for the rendering of images and shapes in a more efficient manner than traditional X11 methods. | Graphics, Rendering, X11 | Computer Science, Computer and Information Sciences | Library | Documentation, Uses, and more |
| libxslt | mcc | N/A | libxslt is the XSLT C library developed for the GNOME project. XSLT itself is a an XML language to define transformation for XML. Libxslt is based on libxml2 the XML C library developed for the GNOME project. Description Source: http://xmlsoft.org/libxslt/index.html |
Xml, Xslt, C Library, Transformation | Software Engineering, Computer & Information Sciences | Data Processing | Documentation, Uses, and more |
| libxsmm | mcc | N/A | libxsmm is a library for specialized dense and sparse matrix operations as well as for deep learning primitives such as small convolutions. The library is targeting Intel Architecture with Intel SSE, Intel AVX, Intel AVX2, Intel AVX‑512 (with VNNI and Bfloat16), and Intel AMX (Advanced Matrix Extensions) supported by future Intel processor code-named Sapphire Rapids. Description Source: https://github.com/libxsmm/libxsmm |
Library, Matrix Multiplication, Intel Cpus | Computer Science, Computer & Information Sciences | Optimization Library | Documentation, Uses, and more |
| libxt | mcc, lcc | N/A | Libxt is a library offering extensions for automating compilation, dependency tracking, and execution of computational tasks within HPC environments. | Library, Hpc, Automation, Compilation, Dependency Tracking, Parallel Processing | Engineering & Technology | Library | Documentation, Uses, and more |
| libxtst | mcc | N/A | Documentation, Uses, and more | ||||
| libxxf86vm | mcc | N/A | libxxf86vm is a library that provides an interface to the XFree86 Video Mode Extension, allowing applications to query and set video modes. | X11, Graphics, Video | Software Engineering, Computer Science | Library | Documentation, Uses, and more |
| liftoff | mcc | 1.0 | Liftoff is an annotation lift-over tool that transfers gene models from a well-annotated reference assembly onto a new target assembly of the same or a closely related species, an alternative to coordinate-based lift-over that instead aligns the actual gene sequences. Rather than relying on a chain file, it extracts each reference gene's sequence, aligns it to the target genome with minimap2, and then finds the placement that best preserves the gene's exon/intron structure, resolving overlaps and paralog mappings. Inputs are the target genome FASTA, the reference genome FASTA, and the reference GFF/GTF annotation; output is a new GFF3 of lifted features plus a report of any genes that failed to map. Version 1.6.3 is widely used to quickly annotate newly assembled genomes and to compare annotations across assemblies. | Build System, Ios Development, Macos Development, Software Development, Apple Platforms | Engineering & Technology | Build Automation | Documentation, Uses, and more |
| liftover | mcc | 1.0 | UCSC liftOver (this build is version 482) converts genomic coordinates and feature annotations between different assembly versions of the same species — for example remapping intervals from human GRCh37/hg19 to GRCh38/hg38, or between species — using precomputed "chain" files that describe the alignment between the two assemblies. It preserves feature identity while translating positions, and reports records that fail to map (because they fall in deleted, split, or ambiguously duplicated regions) to a separate unmapped file. BED is the most common input, but it can also handle other position formats such as GFF via an option. Chain files are downloaded from the UCSC Genome Browser downloads. It is an essential utility whenever coordinate data (variants, peaks, gene models) produced against one reference must be reconciled with another. | genomics, bioinformatics, coordinate conversion, genome assembly | Bioinformatics, Genomics | Command-line tool | Documentation, Uses, and more |
| ligpargen | lcc | N/A | LigParGen is a molecular-mechanics parameterization tool from William L. Jorgensen's group (Yale) that generates OPLS-AA (1.14*CM1A / CM1A-LBCC) force-field parameters, charges, and topologies for small organic molecules and ligands. It accepts SMILES, MOL, or PDB input and emits ready-to-use input files for GROMACS, LAMMPS, NAMD/CHARMM, OpenMM, Tinker, Desmond, BOSS, Q, and others, letting computational chemists prepare drug-like molecules for MD and free-energy simulations. On LCC it is installed as the conda environment LigParGen-2.1 (module ccs/conda/LigParGen-2.1) with the LigParGen driver on PATH. | ligand parameter generation, molecular dynamics, bioinformatics | Molecular Modeling, Computational Chemistry | Standalone | Documentation, Uses, and more |
| likwid | mcc, lcc | N/A | Performance tools for the Linux console | Performance Analysis, Hpc, Performance Counter Profiling, Hardware Performance Counters | Performance Evaluation & Benchmarking, Engineering & Technology | Tool | Documentation, Uses, and more |
| lima | mcc, lcc | 1.0 | lima, version 2.0.0, is PacBio's official demultiplexer for identifying, clipping, and separating barcodes and primer/adapter sequences in SMRT sequencing data, and is the standard first step in HiFi, Iso-Seq, and amplicon workflows before downstream analysis. It scans each read for the supplied barcode/primer sequences, assigns reads to samples, trims the identified adapter regions, and enforces symmetric/asymmetric barcode pairing rules to minimize misassignment. Input is a PacBio BAM (or FASTA) plus a FASTA of barcode/primer sequences; output is per-barcode BAM files (or a single tagged BAM) accompanied by a lima.report and summary of barcode counts and quality. It offers dedicated presets for HiFi barcoding and for Iso-Seq cDNA primer removal. Its stringent scoring and reporting make it the recommended tool for accurate PacBio sample demultiplexing. | Single-Cell RNA-Seq, Data Analysis, Transcriptomics | Biological Sciences | Bioinformatics Tool | Documentation, Uses, and more |
| linux-perf | mcc | N/A | perf v6.3.10 -- perf is a powerful Linux tool for profiling and benchmarking system and application performance by monitoring hardware and software events. | performance, monitoring, profiling, Linux | Computer Science | Performance Analysis Tool | Documentation, Uses, and more |
| llm-launcher | lcc, ecc | 1.0 | LLM-Launcher is an Open OnDemand (OOD) interactive application registered on the ECC cluster that lets users start a large-language-model serving session on GPU compute nodes directly from their web browser, without manually writing SLURM batch scripts. As an OOD interactive app it presents a web form for choosing job resources (GPU type/count, number of cores, memory, and walltime), submits the backing batch job to the scheduler, and then gives the user a connection back to the running LLM service (for example an inference endpoint or notebook front-end) once the job starts. In the Software Discovery Service this entry is a catalog-only stub rather than a built container image: it exists to make the browser-launchable app discoverable and documented alongside the cluster's other software. Researchers use it to spin up on-demand model inference on ECC GPUs for prototyping, prompting, and evaluation work. | Documentation, Uses, and more | |||
| llm-research | lcc | 1.0 | The llm-research environment is a curated Python 3.11 / CUDA 12.4 GPU software stack assembled for large-language-model security research, specifically adversarial-attack and jailbreak studies. It combines the deep-learning foundation (PyTorch with torchvision/torchaudio built for CUDA 12.4) with the HuggingFace ecosystem — Transformers for loading and running open models, Accelerate for multi-GPU/mixed-precision execution, Datasets for benchmark data, and bitsandbytes for 4/8-bit quantized inference of large models on limited GPU memory. On top of this it bundles attack and evaluation libraries: nanoGCG (an efficient implementation of the Greedy Coordinate Gradient adversarial-suffix attack), JailbreakBench (a standardized jailbreak benchmark and dataset), and FastChat (fschat, for serving and conversation templates), together with SpaCy for NLP preprocessing. For evaluating against and orchestrating both hosted and local models it includes API clients for OpenAI and Anthropic plus LiteLLM (pinned <1.49) as a unified multi-provider interface, and Weights & Biases for experiment tracking, with JupyterLab for interactive work. It is intended for researchers red-teaming and benchmarking model robustness on GPU nodes. | Documentation, Uses, and more | |||
| llvm | mcc, lcc | N/A | The LLVM Project is a collection of modular and reusable compiler and toolchain technologies. Despite its name, LLVM has little to do with traditional virtual machines. Description Source: https://llvm.org/ |
Compiler, Toolchain | Software Engineering, Computer & Information Sciences | Compiler | Documentation, Uses, and more |
| llvm5 | lcc | N/A | LLVM Compiler Infrastructure | Compiler, Programming Language, Optimization, Toolchain | Software Engineering, Computer Science | Compiler Infrastructure | Documentation, Uses, and more |
| lmod | ecc | N/A | Lmod is a Lua based module system that easily handles the MODULEPATH Hierarchical problem. Environment Modules provide a convenient way to dynamically change the users’ environment through modulefiles. This includes easily adding or removing directories to the PATH environment variable. Description Source: https://lmod.readthedocs.io/en/latest/ |
Module Management, Software Environment, Environment Module System | Infrastructure & Instrumentation, Engineering & Technology | Utility | Documentation, Uses, and more |
| longstitch | mcc | 1.0 | LongStitch (version 1.0.5) is a Snakemake-based pipeline for correcting and scaffolding draft genome assemblies using long reads from PacBio or Oxford Nanopore platforms. It chains together three ABySS-family tools: Tigmint-long, which detects and cuts assembly errors by mapping long reads back to the draft; ntLink, a minimizer-based scaffolder that also gap-fills and can be run iteratively; and optionally ARKS/ARCS-long for additional long-read scaffolding. The result is a more contiguous, error-corrected assembly with improved N50 and fewer misassemblies. Inputs are a draft assembly FASTA and a long-read FASTA/FASTQ file, and the pipeline can select among the tigmint-ntLink, ntLink-arks, or full tigmint-ntLink-arks target combinations. It is intended as a lightweight polishing step in long-read genome assembly workflows. | Documentation, Uses, and more | |||
| ltr_retriever | mcc | 1.0 | LTR_retriever is a program for the accurate identification and annotation of long terminal repeat (LTR) retrotransposons in genome assemblies, delivered here as part of the Dfam TE Tools (TETools) v1.9.5 curated transposable-element annotation environment packaged from the upstream dfam/tetools:1.95 Docker image. LTR_retriever takes the candidate LTR predictions produced by structure-based finders such as LTRharvest and LTR_FINDER and rigorously filters and refines them — recovering intact elements, resolving nested and truncated insertions, and greatly reducing false positives — to build a high-quality, non-redundant LTR retrotransposon library. That library can then feed genome-wide repeat masking and, notably, the calculation of the LTR Assembly Index (LAI), a widely used measure of assembly contiguity in repeat regions. The surrounding TETools image also bundles the standard Dfam/RepeatMasker/RepeatModeler stack (RepeatMasker, RepeatModeler, RMBlast, etc.), so the container supports a complete TE-discovery and annotation workflow. Inputs are a genome FASTA plus candidate LTR coordinates; outputs are refined LTR annotations and a curated repeat library. | bioinformatics, genomics, LTR retrotransposons, sequence analysis | Computational Biology, Genomics | Bioinformatics Tool | Documentation, Uses, and more |
| lua | ecc | N/A | Lua is a powerful, fast, lightweight, embeddable scripting language. Lua combines simple procedural syntax with powerful data description constructs based on associative arrays and extensible semantics. Lua is dynamically typed, runs by interpreting bytecode for a register-based virtual machine, and has automatic memory management with incremental garbage collection, making it ideal for configuration, scripting, and rapid prototyping. | Programming Language, Scripting Language, Embedded Systems | Software Engineering, Computer & Information Sciences | Programming Language | Documentation, Uses, and more |
| lua-luafilesystem | ecc | N/A | LuaFileSystem is a Lua library that provides a portable way to access the underlying filesystem. It allows users to perform operations such as file and directory manipulation, path manipulation, and more. | Lua, Filesystem, Library, Cross-platform | Software Engineering, Other Computer and Information Sciences | Library | Documentation, Uses, and more |
| lua-luaposix | ecc | N/A | LuaPosix is a Lua binding for the POSIX API, providing access to operating system functionalities such as file manipulation, process control, and threading. | Lua, POSIX, Bindings, Operating System | Software Engineering, Other Computer and Information Sciences | Library | Documentation, Uses, and more |
| lumerical | lcc | N/A | Lumerical software. | photonic design, electromagnetic simulation, optical analysis, engineering | Engineering, Electrical, Electronic, and Information Engineering | Commercial | Documentation, Uses, and more |
| lz4 | mcc, lcc, ecc | N/A | LZ4 is lossless compression algorithm, providing compression speed > 500 MB/s per core (>0.15 Bytes/cycle). It features an extremely fast decoder, with speed in multiple GB/s per core (~1 Byte/cycle). A high compression derivative, called LZ4_HC, is available, trading customizable CPU time for compression ratio. Description Source: https://lz4.org/ |
Compression, Data Compression, Lossless Compression | Computer Science, Computer & Information Sciences | Utility | Documentation, Uses, and more |
| lzo | ecc | N/A | LZO offers pretty fast compression and extremely fast decompression. MiniLZO is a very lightweight subset of the LZO library. Description Source: https://github.com/conda-forge/lzo-feedstock/blob/main/recipe/meta.yaml | Data Compression, Algorithm, High-Speed | Computer Science, Computer & Information Sciences | Data Processing | Documentation, Uses, and more |
| m4 | mcc, lcc, ecc | N/A | GNU M4 is an implementation of the traditional Unix macro processor. It is mostly SVR4 compatible although it has some extensions (for example, handling more than 9 positional parameters to macros). GNU M4 also has built-in functions for including files, running shell commands, doing arithmetic, etc. Description Source: https://www.gnu.org/software/m4/ |
Macro Processor, Text Processing, Software Development, Build Process | Software Engineering, Computer & Information Sciences | Compiler | Documentation, Uses, and more |
| maaslin2 | lcc | 1.0 | MaAsLin2 (Microbiome Multivariable Associations with Linear Models) is an R/Bioconductor tool from the Huttenhower lab for discovering statistically significant associations between microbial community features and sample metadata in population-scale microbiome studies. It fits per-feature multivariable general linear models relating feature abundances (taxa, gene families, pathways, or other omics measurements) to multiple covariates simultaneously, supporting fixed and random effects for repeated-measures/longitudinal designs, a range of normalization (TSS, CLR, etc.) and transformation (log, arcsin-square-root) choices, several model families (linear, negative-binomial, zero-inflated), and multiple-testing correction. Inputs are a feature-abundance table and a metadata table; outputs include a table of significant associations with coefficients, p-values, and q-values, plus diagnostic and effect-size visualizations. Version 0.99.12 is provided as its own conda environment in a Rocky 8 + Miniconda multi-tool bioinformatics container. | bioinformatics, microbiome, statistics, data analysis | Bioinformatics, Microbial Ecology | Statistical Analysis Tool | Documentation, Uses, and more |
| macs | lcc | 1.0 | MACS2 (Model-based Analysis of ChIP-Seq), version 2.2.7.1, is the standard tool for identifying regions of read enrichment ("peaks") from ChIP-seq and related experiments such as ATAC-seq. It empirically models the fragment-length shift from the bimodal read distribution, then uses a local dynamic Poisson background to call statistically significant peaks while controlling false discovery rate, and supports both narrow-peak calling (for transcription factors) and broad-peak calling (for diffuse histone marks). Its callpeak subcommand takes treatment and (optionally) control alignment files in BED/BAM/other formats and outputs narrowPeak/broadPeak BED files, summit coordinates, and bedGraph pileup and fold-enrichment tracks for genome-browser visualization; additional subcommands (bdgcmp, bdgpeakcall, pileup) support custom track and signal analysis. It is a mainstay of epigenomics and gene-regulation studies. | ChIP-Seq, Bioinformatics, Genomics, Peak Calling | Bioinformatics, Genomics | Analysis Tool | Documentation, Uses, and more |
| macs2 | mcc, lcc | 1.0 | MACS2 (Model-based Analysis of ChIP-Seq, version 2) is the standard peak-caller for identifying regions of genomic enrichment from ChIP-seq, ATAC-seq, and related pull-down or open-chromatin assays. Its main subcommand, callpeak, takes aligned reads (BED/BAM, single- or paired-end) for a treatment sample and an optional control/input sample, empirically models the fragment-length shift from the read distribution, estimates local background using a dynamic Poisson model (lambda), and calls significantly enriched peaks with associated fold-enrichment and q-values, outputting narrowPeak/broadPeak BED files, peak summits, and coverage tracks (bedGraph). It supports both narrow (transcription-factor) and broad (histone-mark) peak modes, effective-genome-size specification, and duplicate handling. Other subcommands (bdgcmp, bdgdiff, bdgpeakcall, pileup, etc.) support signal-track generation and differential peak analysis. This build is version 2.2.7.1 from Bioconda. | Chip-Seq Analysis, Transcription Factor Binding Sites, Histone Modification, Bioinformatics | Biological Sciences | Command-Line Tool | Documentation, Uses, and more |
| mafft | mcc, lcc | 1.0 | MAFFT (version 7.525) is a fast and widely used program for multiple sequence alignment of nucleotide or amino-acid sequences. It offers a range of strategies that trade off speed against accuracy: progressive methods (FFT-NS-1, FFT-NS-2) for very large datasets, iterative-refinement methods (FFT-NS-i, L-INS-i, G-INS-i, E-INS-i) that incorporate local or global pairwise consistency information for higher accuracy on smaller or difficult sets, and specialized modes for aligning sequences to an existing alignment or adding fragmentary sequences. It uses fast Fourier transform to identify homologous regions efficiently and scales to thousands of sequences. Input is a FASTA file and output is an alignment in FASTA (or other) format, and its convenient automatic mode selects an appropriate algorithm based on data size. MAFFT is a standard first step in phylogenetics, comparative genomics, and profile/HMM building. | Bioinformatics, Sequence Alignment, Phylogenetics | Bioinformatics, Biological Sciences | Bioinformatics Tool | Documentation, Uses, and more |
| make | lcc | N/A | Make is an open-source, tool designed to build, test and package software. | Build Automation, Software Development, Programming | Computer Science, Engineering & Technology | Utility | Documentation, Uses, and more |
| make_tracks_file | mcc | 1.0 | make_tracks_file is a convenience command shipped with pyGenomeTracks (version 3.9 from bioconda) that auto-generates a starter INI configuration describing a set of genomic tracks, saving users from writing the tracks file by hand. Given a list of input files (e.g., BED, bigWig, bedGraph, Hi-C matrices, GTF), it produces an editable .ini with one section per track and sensible default styling, which the user can then customize (colors, height, labels, display type) before rendering. It is used together with the main pyGenomeTracks command, which reads the INI plus a genomic region and outputs a publication-quality stacked track figure (PDF/PNG/SVG). The typical two-step workflow first auto-generates the tracks configuration from a set of input files and then renders a chosen genomic region into an image. It is aimed at genomics researchers producing browser-style multi-track visualizations for figures. | Documentation, Uses, and more | |||
| maker | mcc, lcc | 1.0 | MAKER (BioContainers image 3.01.03) is a portable, extensible genome-annotation pipeline that integrates multiple lines of evidence to produce, refine, and update structural gene annotations, and is especially popular for emerging/non-model organisms. It combines ab initio gene predictors (such as SNAP, Augustus, and GeneMark) with aligned empirical evidence — EST/transcript and protein homology alignments (via BLAST and Exonerate) and repeat masking (RepeatMasker) — to generate consensus gene models with quality scores (annotation edit distance, AED). It supports iterative training, where initial annotations bootstrap and retrain the predictors for improved subsequent rounds, and can incorporate or update existing GFF3 annotations. Configuration is driven by control files (maker_opts.ctl, maker_bopts.ctl, maker_exe.ctl) that point to the genome, evidence files, and executables; MAKER runs (optionally under MPI) and outputs GFF3 gene models plus predicted transcript and protein FASTA, gathered with tools like gff3_merge and fasta_merge. This container wraps the upstream image, adding only the cluster mount points. | Genome Annotation, Bioinformatics | Genomics, Biological Sciences | Annotation Tool | Documentation, Uses, and more |
| maple | lcc | N/A | maple tool designed to package software. | Mathematics, Symbolic Computation, Numerical Analysis, Data Visualization | Computational Mathematics, Applied Mathematics | Commercial | Documentation, Uses, and more |
| mapsplice | lcc | N/A | MapSplice is a bioinformatics application for the alignment of RNA-seq reads to a reference genome with a focus on accurate detection of splice junctions. It is widely used in transcriptomics to map spliced reads and discover both canonical and novel exon-exon junctions without relying on gene annotations. Built here with the GNU8 compiler toolchain and OpenMPI3 stack, it is typically used in gene-expression and alternative-splicing analysis pipelines. | RNA-Seq, Splice Junction Detection, Bioinformatics, Genomics | Bioinformatics, Genomics | Bioinformatics Tool | Documentation, Uses, and more |
| marty | lcc | 1.0 | MARTY (Modern ARtificial Theoretical phYsicist) is a C++ framework for automated symbolic computation in high-energy particle physics, aimed at beyond-the-Standard-Model (BSM) model building. Given a Lagrangian defined through its companion symbolic computer-algebra library (CSL), MARTY can generate the Feynman rules of a theory, compute tree-level and one-loop amplitudes and squared amplitudes analytically, and derive Wilson coefficients for effective field theories used in flavor physics and other precision observables. It handles group theory, gauge symmetries, and the algebra of gamma matrices and color/flavor structures symbolically, and can output results as generated C++ code (a numerical library) that the user compiles to evaluate cross sections, decay rates, or Wilson coefficients numerically. The typical workflow is to write a C++ program describing the model and the requested computation, link against the MARTY/CSL libraries, and compile and run it. This deployment builds MARTY version 1.3 from the upstream source tarball on an Ubuntu 18.04 base. | Documentation, Uses, and more | |||
| mashmap | mcc, lcc | 1.0 | MashMap is a fast approximate long-sequence aligner that computes the approximate local alignment boundaries and identity between long DNA sequences without producing base-level alignments, making it well suited to comparing genome assemblies, mapping long reads, and whole-genome dot-plot analyses. It uses MinHash sketching and a minimizer-based approach to estimate the Jaccard similarity and hence sequence identity over windows, reporting mapped regions above a user-set identity threshold and minimum segment length. Inputs are FASTA query and reference sequences, and output is a PAF-like table of mapped coordinates and estimated identity, which pairs well with visualization tools for synteny/dot plots. This is version 3.1.3, which uses a minimal chaining step for improved region boundaries. | Mapping, Alignment, Genomics, Bioinformatics | Genomics, Biological Sciences | Bioinformatics | Documentation, Uses, and more |
| mashtree | lcc | 1.0 | Mashtree builds phylogenetic-style trees from whole-genome assemblies very rapidly by using Mash (MinHash) sketches to estimate pairwise genomic distances rather than performing a full multiple-sequence alignment. Each input genome (FASTA, or raw reads) is reduced to a compact sketch of representative k-mers; Mash then estimates the Jaccard-based distance between every pair, and those distances are assembled into a distance matrix from which a neighbor-joining tree is produced. Because it avoids alignment and gene-by-gene comparison, it can relate hundreds or thousands of genomes in minutes, making it ideal for quick clustering, outbreak screening, and dereplication—though the resulting tree is a distance dendrogram, not a substitution-model phylogeny. It offers options to add bootstrap-like confidence values. Version 1.2.0 is popular in microbial genomics for fast, approximate relatedness assessment. | Phylogenetic Trees, Genomic Similarity, Mash Distances | Genomics, Biological Sciences | Tool | Documentation, Uses, and more |
| masurca | mcc, lcc | 1.0 | MaSuRCA (Maryland Super-Read Cabog Assembler) is a whole-genome assembler that combines the strengths of de Bruijn graph and overlap-layout-consensus (OLC) approaches by first transforming large numbers of short reads into a much smaller set of longer, highly accurate 'super-reads', which are then assembled. It supports assembly of Illumina short reads alone as well as hybrid assembly incorporating PacBio or Oxford Nanopore long reads, and it bundles the polishing/consensus and scaffolding steps needed to produce a final assembly. It is driven by a configuration file that lists the read libraries (with insert sizes) and parameters; from that configuration it generates an assembly script that executes the full pipeline. This is version 4.0.8. | Genome Assembly, Bioinformatics, Computational Biology | Bioinformatics, Biological Sciences | Bioinformatics Tool | Documentation, Uses, and more |
| matlab | mcc, lcc | 1.0 | Matlab | Programming Language, Technical Computing, Visualization, Mathematical Notation | Applied Mathematics, Mathematics | Computational Software | Documentation, Uses, and more |
| matplotlib | lcc | N/A | Matplotlib is a comprehensive library for creating static, animated, and interactive visualizations in Python. Matplotlib makes easy things easy and hard things possible. Description Source: https://matplotlib.org/ |
Python Library, Data Visualization, Plotting Library | Computer Science, Computer & Information Sciences | Plotting & Graphing Tool | Documentation, Uses, and more |
| mbg | mcc | 1.0 | Minimizer based sparse de Bruijn graph builder, the graph construction engine that ribotin calls on the reads it recruits and the same tool that builds the initial HiFi graph inside Verkko. It assembles accurate long reads into a sparse de Bruijn graph in GFA form using minimizer based node selection, which keeps memory and run time low enough for repetitive regions such as rDNA arrays and satellites where ordinary overlap assembly struggles. It is included here because ribotin invokes it as an external program rather than linking it, and refuses to run without it on PATH. It is also useful on its own for producing an assembly graph from HiFi or duplex reads, including the partial reference morphs that non human ribotin runs need as input. | Documentation, Uses, and more | |||
| mcclintock | mcc | 1.0 | McClintock is a meta-pipeline for detecting transposable-element (TE) insertions from short-read whole-genome sequencing that runs and integrates the output of many independent TE-detection component methods, since no single caller reliably captures all insertion types. Orchestrated with Snakemake and per-method conda/mamba environments, it wraps tools such as ngs_te_mapper, RelocaTE, TEMP/TEMP2, RetroSeq, PoPoolationTE, TEbreak, TEFLoN and others, running them from a common set of inputs—a reference genome, a TE library/consensus sequences, an optional TE annotation, and the sequencing reads—and producing standardized, comparable BED files of predicted reference and non-reference TE insertions for each method. This lets researchers compare methods or take a consensus, improving confidence in TE-polymorphism calls. It is driven through a Python driver script pointing at the config and inputs, and is used in evolutionary and population genomics of mobile-element variation. | Documentation, Uses, and more | |||
| mcl | lcc | 1.0 | MCL (version 14.137) is the reference implementation of the Markov Cluster algorithm, a fast, scalable, unsupervised method for clustering nodes in a weighted graph or network. It works by simulating stochastic flow (random walks) on the graph through alternating "expansion" and "inflation" operations on the transition matrix, where the inflation parameter tunes cluster granularity — higher inflation yields more, smaller clusters. Because it discovers natural groupings without requiring the number of clusters to be specified in advance, MCL is a popular choice in bioinformatics, most notably for clustering proteins into families from all-versus-all sequence-similarity graphs (as used by OrthoMCL and TribeMCL for orthology/gene-family inference). Input is an abstract weighted graph, typically supplied as a tab-separated list of edges with weights (label mode) or an MCL matrix; output is a partition assigning each node to a cluster. It is also used generally for network and community-detection tasks. | Clustering, Network Analysis, Graph Algorithms | Computational Biology, Computer & Information Sciences | Clustering Algorithm | Documentation, Uses, and more |
| mcscanx | mcc | 1.0 | MCScanX detects and visualizes syntenic (collinear) blocks of genes within a single genome or between genomes, which is the basis for studying whole-genome and segmental duplications, gene-family expansion, and genome evolution. It takes the results of an all-versus-all sequence-similarity search (typically BLASTp) plus a gene-position (GFF-derived) file, then chains homologous gene pairs that occur in conserved order into collinear blocks using a dynamic-programming algorithm, classifying genes by duplication origin (singleton, dispersed, proximal, tandem, and whole-genome/segmental). It ships downstream utilities to compute Ka/Ks on collinear pairs and to render dual synteny plots, circle plots, dot plots, and bar charts of the detected blocks. A typical workflow prepares a BLAST file and a matching GFF file, runs MCScanX to produce a collinearity file, then applies the plotting tools to that result. Version 1.0.0 is provided as a conda app within a multi-app Miniconda container focused on comparative/structural genomics and population structure (alongside NGenomeSyn, plotsr, SyRI, SYNY, CLUMPP, KMC, and vg). | Gene Synteny, Genomic Analysis, Syntenic Relationships | Bioinformatics, Biological Sciences | Tool | Documentation, Uses, and more |
| mdanalysis | lcc | 1.0 | MDAnalysis is an object-oriented Python library for analyzing trajectories from molecular dynamics (MD) simulations, providing a common interface to the output formats of essentially all major MD engines (GROMACS, AMBER, NAMD, CHARMM, LAMMPS, and others). It loads a topology plus one or more trajectory files into a Universe object, exposing a powerful atom-selection language (for example selecting protein alpha-carbons or atoms within a given distance of a ligand residue) and NumPy-backed access to coordinates over time, enabling custom and built-in analyses such as RMSD/RMSF, radius of gyration, hydrogen bonds, radial distribution functions, density profiles, contacts, and diffusion. It is used programmatically in Python scripts and notebooks, iterating frame-by-frame over a trajectory or applying the analysis-module classes, and it interoperates cleanly with the scientific Python stack for further statistics and plotting. This is conda-forge mdanalysis version 1.0.0. | Molecular Dynamics, Simulation Analysis, Python, Scientific Computing | Molecular Dynamics, Biochemistry and Molecular Biology | Library | Documentation, Uses, and more |
| mdrun | mcc | 1.0 | Molecular dynamics engine of GROMACS 4.5.4, here built with the SMOG structure-based model extension, used to integrate the equations of motion for biomolecular simulations and in particular for structure-based (Go-like) models of protein folding. It evaluates the Gaussian contact potentials added by the extension, reporting their contribution as a separate Gaussian term in the energy log alongside the usual bonded and nonbonded terms, and it runs ordinary GROMACS 4.5 simulations unchanged. This binary uses shared-memory thread parallelism within a single node, chosen by a thread count at run time, while a separate message-passing binary suffixed _mpi covers multi-node runs. Single and double precision engines are both installed, the latter suffixed _d, and double precision is the safer choice for long runs where energy conservation is being monitored. | Documentation, Uses, and more | |||
| mdrun_mpi | mcc | 1.0 | Message-passing build of the GROMACS 4.5.4 molecular dynamics engine, for runs that need more cores than a single node provides. It is launched by the host MPI, one process per rank, and distributes the simulation across ranks by domain decomposition rather than by shared-memory threads, so the rank count alone sets the parallelism. Like the rest of this build it understands the SMOG Gaussian contact potentials, and it is equally usable for ordinary GROMACS 4.5 simulations that have nothing to do with structure-based models. A double precision counterpart suffixed _mpi_d is installed alongside it, and the thread-based engines remain the better choice for work that fits on one node. | Documentation, Uses, and more | |||
| mdtraj | lcc | 1.0 | MDTraj is a Python library for reading, writing, and analyzing molecular dynamics (MD) trajectories, version 1.9.5. It provides fast, memory-efficient loading of a wide range of trajectory and topology formats — including DCD, XTC/TRR (GROMACS), NetCDF (AMBER), HDF5, LAMMPS, and PDB — into a unified Trajectory object backed by NumPy arrays, and interoperates smoothly with the scientific Python stack. Its analysis routines cover geometric measurements (distances, angles, dihedrals), RMSD and RMSF with fast alignment, radius of gyration, hydrogen-bond and secondary-structure (DSSP) assignment, solvent-accessible surface area, native contacts, and more, plus format conversion between simulation packages. It is imported as a Python module and used to load trajectories with an associated topology and run analyses such as RMSD. This environment additionally installs networkx 2.5.1 and MDAnalysis, a complementary trajectory-analysis library, giving users two major Python MD-analysis toolkits side by side. | Molecular Dynamics, Trajectory Analysis, Python, Bioinformatics | Molecular Dynamics, Biochemistry and Molecular Biology | Library | Documentation, Uses, and more |
| med | lcc | N/A | Documentation, Uses, and more | ||||
| medaka | mcc | 1.0 | medaka is Oxford Nanopore Technologies' tool for producing high-accuracy consensus sequences and variant calls from nanopore reads using neural networks trained to correct the systematic error patterns of nanopore basecalling. Its most common use, the medaka_consensus workflow, polishes a draft assembly: nanopore reads are aligned to the draft, medaka's recurrent neural network model processes the pileup to predict a corrected consensus, and it outputs a polished assembly FASTA, substantially reducing residual indel and substitution errors. It also provides variant-calling workflows (medaka_variant) for calling SNPs/indels against a reference from nanopore data, and supports selecting basecaller-matched models so the neural network matches the chemistry and basecaller used. Version 1.7.2 is provided as a conda app within a multi-app Miniconda container that also bundles genome alignment/assembly, expression clustering, visualization, and file-transfer tools (SibeliaZ, lftp, ParaView, clust, Verkko, SPAdes, and Flye). | Bioinformatics, Ngs Data Analysis, Nanopore Sequencing, Consensus Sequence Calling | Bioinformatics, Biological Sciences | Ngs Data Analysis Tool | Documentation, Uses, and more |
| megahit | lcc | 1.0 | MEGAHIT (version 1.2.9) is an ultra-fast and memory-efficient de novo assembler built specifically for large and complex metagenomic short-read datasets, though it also assembles single-genome and transcriptome data. It uses succinct de Bruijn graphs and a multiple-k-mer strategy, iterating over a range of k values (with mercy k-mers to recover low-coverage regions) to balance sensitivity for rare species with contiguity for abundant ones. Because of its low memory footprint it can assemble very large soil or gut metagenomes on modest hardware. Inputs are paired-end and/or single-end FASTQ/FASTA reads and the output is a contigs FASTA, with the user able to specify the list of k-mer sizes to iterate over. It is one of the most popular first-step assemblers in metagenomics binning and MAG-recovery pipelines. | Metagenome Assembly, Bioinformatics, Sequencing Data, Snp-Aware, Draft Assemblies | Bioinformatics, Biological Sciences | Assembly Software | Documentation, Uses, and more |
| megahit-1.2.9 | mcc | 1.0 | MEGAHIT is an ultra-fast and memory-efficient de novo assembler specialized for large and complex metagenomics short-read datasets, though it also handles single-genome and transcriptome data. It builds a succinct de Bruijn graph and iterates assembly over multiple k-mer sizes to balance sensitivity for low-abundance genomes with contiguity for high-abundance ones, allowing large metagenomes to be assembled within modest memory footprints. It takes paired-end and/or single-end FASTQ/FASTA reads and outputs contigs in FASTA along with an assembly log, with optional tuning of the k-mer list and minimum contig length. This is version 1.2.9. | metagenomics, assembly, bioinformatics | Bioinformatics, Biological Sciences | Command-line tool | Documentation, Uses, and more |
| megan | lcc | 1.0 | MEGAN (MEtaGenome ANalyzer) is an interactive tool for the taxonomic, functional, and comparative analysis of metagenomic (and metatranscriptomic) sequencing data. Its typical workflow starts from the results of a sequence-similarity search (e.g. DIAMOND or BLAST of reads against NR/a protein database); the accompanying daa2rma/blast2rma converters ingest those alignments and MEGAN assigns each read to a taxon using the naive lowest-common-ancestor (LCA) algorithm, producing an interactive taxonomic tree that can be explored, collapsed, and compared across samples. Beyond taxonomy it performs functional classification by mapping reads to functional hierarchies such as InterPro2GO, SEED, eggNOG, and KEGG, and supports comparative multi-sample analyses with rarefaction, PCoA, and clustering. It provides both a rich GUI and command-line tools for batch processing. It is a long-standing choice for making sense of shotgun metagenomes. This build is version 6.24.20 from Bioconda (the free MEGAN Community Edition). | Metagenomics, Microbiome, Bioinformatics | Bioinformatics, Biological Sciences | Bioinformatics Tool | Documentation, Uses, and more |
| megax | lcc | 1.0 | MEGA (Molecular Evolutionary Genetics Analysis), the X (version 10) release, is a comprehensive application for comparative and phylogenetic analysis of DNA and protein sequences. It provides an integrated environment for aligning sequences (via bundled MUSCLE and ClustalW), estimating evolutionary distances under a wide range of nucleotide and amino-acid substitution models, inferring phylogenetic trees by several methods (Neighbor-Joining, Maximum Parsimony, Maximum Likelihood, UPGMA, and Minimum Evolution) with bootstrap support, and testing evolutionary hypotheses (selection tests such as dN/dS, molecular-clock tests, and ancestral-sequence reconstruction). MEGA X is well known for its user-friendly graphical interface, but it also ships a command-line 'MEGA-CC' (computational core) driven by saved analysis-option (.mao) files for batch/scripted and HPC use, which is the relevant mode in a container/cluster context. It is a longstanding teaching and research tool in molecular evolution. This deployment installs MEGA X version 10.1.8 from the vendor RPM (megasoftware.net). | Documentation, Uses, and more | |||
| merqury | mcc | 1.0 | Merqury is a reference-free tool for evaluating the quality and completeness of genome assemblies using k-mer spectra derived from the raw sequencing reads rather than an external reference genome. By counting k-mers in the high-accuracy reads (using Meryl) and comparing them to the k-mers present in the assembly, Merqury estimates a consensus quality value (QV, a phred-scaled base-accuracy estimate), k-mer completeness (the fraction of reliable read k-mers represented in the assembly), and produces copy-number/spectra-cn plots that reveal collapsed or duplicated regions. For trio or phased assemblies it also assesses haplotype phasing and switch errors using parental k-mer sets (hap-mers). The standard workflow first builds a meryl k-mer database from the reads and then runs Merqury against that database and the assembly, generating QV/completeness tables and diagnostic plots. It is a standard assembly-QC step, especially for long-read and HiFi assemblies. This build is version 1.3 from Bioconda and pulls in meryl, samtools, bedtools, R, and a JDK. | Computational Tool, Genomics, Bioinformatics | Biological Sciences | Bioinformatics | Documentation, Uses, and more |
| meryl | lcc | N/A | Meryl 1.0 (marbl/meryl) is a genomics tool for counting, indexing, and performing set operations on k-mers in DNA sequence data. It ships the meryl, meryl-import, meryl-lookup, and sequence binaries, and underlies assemblers and assembly-QC tools such as Canu and Merqury. Genome scientists use it for k-mer spectrum analysis, genome size estimation, and assembly evaluation. | Bioinformatics, Genomics, K-Mer Counting | Bioinformatics, Biological Sciences | Bioinformatics Tool | Documentation, Uses, and more |
| mesa | mcc, lcc | N/A | The Mesa project began as an open-source implementation of the OpenGL specification - a system for rendering interactive 3D graphics. Over the years the project has grown to implement more graphics APIs, including OpenGL ES, OpenCL, OpenMAX, VDPAU, VA-API, Vulkan and EGL. Description Source: https://docs.mesa3d.org/index.html |
Agent-Based Modeling, Complex Systems, Simulation, Computational Modeling | Other Computer & Information Sciences, Computer & Information Sciences | Framework | Documentation, Uses, and more |
| mesa-glu | mcc, lcc | N/A | mesa-gl is an open-source implementation of the OpenGL specification, a system for rendering interactive 3D graphics, and a widely used and versatile graphics library. mesa-glu is the OpenGL Utility Library (GLU) component of Mesa, providing useful functions for building OpenGL applications. | Graphics, 3D Graphics, Opengl, Graphics Library | Library | Documentation, Uses, and more | |
| meshio | lcc | N/A | meshio is a Python library for reading and writing mesh files in various formats, making it easier to work with finite element analysis (FEA) and computational fluid dynamics (CFD) data. | mesh, finite element analysis, computational fluid dynamics, Python | Engineering, Mechanical Engineering | Library | Documentation, Uses, and more |
| meshlab | mcc | 1.0 | MeshLab is an open-source system for processing, editing, and inspecting unstructured 3D triangular meshes and point clouds, used in 3D scanning, cultural-heritage digitization, reverse engineering, and scientific visualization. It provides a large library of filters for cleaning (removing duplicate/non-manifold elements, filling holes), remeshing and simplification (quadric edge-collapse decimation), surface reconstruction from point clouds (e.g. Poisson), normal and curvature estimation, alignment/registration of range scans, measurement, and high-quality rendering. It supports many mesh/point-cloud formats (PLY, OBJ, STL, OFF, 3DS, and more) for import, conversion, and export. This build is the Ubuntu 22.04 package and includes the pymeshlab Python bindings, so operations can be performed interactively in the GUI or scripted/batched from Python. | 3D Modeling, Mesh Processing, Open Source | 3D Modeling and Visualization, Computer Science | Desktop Application | Documentation, Uses, and more |
| meson | mcc, lcc, ecc | N/A | Meson is an open source build system meant to be both extremely fast, and, even more importantly, as user friendly as possible. Description Source: https://mesonbuild.com/ |
Build System, Software Development, Automation | Computer Science, Computer & Information Sciences | Compiler | Documentation, Uses, and more |
| metabat2 | mcc | 1.0 | MetaBAT2 (from bioconda, here installed into the anvio-7.1 environment) is a widely used tool for metagenome binning: grouping assembled contigs into draft genome bins (metagenome-assembled genomes, MAGs) that each ideally represent a single microbial population. It clusters contigs using a combination of tetranucleotide-frequency composition signals and per-sample coverage (abundance) profiles, an approach that is largely parameter-free in version 2 and scales to large, complex communities. The typical workflow first derives a coverage depth file from sorted BAMs and then bins the assembly contigs using that depth information, producing one FASTA per recovered bin. Binning quality is then usually assessed with CheckM/BUSCO and the bins refined or dereplicated. As part of an anvi'o-centric container, it fits naturally into metagenomic assembly-to-MAG pipelines. | metagenomics, binning, bioinformatics, sequence analysis | Bioinformatics, Metagenomics | Command-line tool | Documentation, Uses, and more |
| metaphlan | mcc, lcc | 1.0 | MetaPhlAn (Metagenomic Phylogenetic Analysis), version 4.1.0, profiles the taxonomic composition of microbial communities from shotgun metagenomic sequencing without assembly, using a curated database of clade-specific marker genes. By mapping reads (via Bowtie2) to roughly a million unique markers spanning bacteria, archaea, viruses, and eukaryotes, it estimates the relative abundance of taxa down to species and, in the 4.x line, SGB (species-level genome bin) resolution, while also enabling strain-level analysis through its companion StrainPhlAn workflow. Input is FASTQ/FASTA reads; output is a tab-delimited profile of taxa and relative abundances, plus an intermediate SAM/bowtie2out file that can be reused. A companion table-merging utility combines many samples into an abundance matrix for downstream diversity and comparative analysis. It is a standard component of the bioBakery meta'omics suite. | Metagenomics, Microbiome, Microbial Ecology, Bioinformatics | Environmental Biology, Biological Sciences | Tool | Documentation, Uses, and more |
| methpipe | mcc | 1.0 | MethPipe is a computational toolkit for analyzing whole-genome bisulfite-sequencing (WGBS) and related DNA-methylation data, taking mapped bisulfite reads through to methylation-level estimation and higher-order feature detection. From aligned reads it computes per-cytosine methylation levels via its methcounts stage, then identifies biologically meaningful regions: hypomethylated regions/HMRs (often marking regulatory elements) via a hidden Markov model, partially methylated domains (PMDs), allele-specific methylation and hairpin/AMR analysis, and differentially methylated regions (DMRs) between samples. It works primarily with the CpG/CpH methylation-count format and provides utilities for symmetric-CpG merging, coverage handling, and format conversion. Typical steps run the methcounts stage on a sorted read file to get per-site levels, then the hmr stage to call hypomethylated regions. It is a well-established pipeline in epigenomics. This is bioconda methpipe version 5.0.1. | Documentation, Uses, and more | |||
| methyldackel | mcc, lcc | 1.0 | MethylDackel is a lightweight, fast tool for extracting per-base DNA methylation measurements from coordinate-sorted, indexed alignments of bisulfite-sequencing (BS-seq/WGBS/RRBS) data. Its main extract command tallies methylated and unmethylated read counts at each cytosine and reports methylation levels in CpG, and optionally CHG and CHH, contexts, writing bedGraph (or optionally methylKit-, cytosine-report-, or merged-context) output suitable for downstream differential-methylation analysis. It includes an mbias command that computes per-position methylation bias plots to help choose read-trimming boundaries that remove end-of-read artifacts, and options to filter by mapping and base quality, handle CpGs spanning both strands, and account for variant positions. A representative workflow first runs the mbias command to diagnose bias, then runs extract with merged-context output to produce a CpG methylation track. Version 0.5.2 is provided as a conda-based tool within a broad CentOS 8 bioinformatics container. | DNA Methylation, Bisulfite Sequencing, Epigenetics | Genetics, Biological Sciences | Data Analysis Tool | Documentation, Uses, and more |
| metis | mcc, lcc | N/A | Metis development files | Computational Software, Graph Partitioning, Mesh Partitioning, Hypergraph Partitioning | Computer Science, Computer & Information Sciences | Tool | Documentation, Uses, and more |
| mfem | lcc | N/A | Lightweight, general, scalable C++ library for finite element methods | Finite Element Method, C++ Library, Parallel Computing | Applied Mathematics, Engineering & Technology | Library | Documentation, Uses, and more |
| mhm2 | mcc | 1.0 | MetaHipMer2 (MHM2) is a de novo metagenome assembler designed for very large short-read metagenomic datasets on high-performance computing systems. It is built on UPC++ and GASNet, a partitioned-global-address-space (PGAS) programming model that lets the assembler distribute the de Bruijn graph and read data across the memory of many nodes, enabling assembly of terabase-scale metagenomes that exceed the memory of a single machine. It performs iterative k-mer analysis, contig generation, local assembly, and scaffolding, producing metagenome contigs/scaffolds from paired-end reads. It runs as an MPI-style distributed job (launched with the appropriate parallel launcher) taking FASTQ input and emitting assembled sequences with assembly statistics. This container packages MHM2 from the upstream cimendes/mhm2 Docker image and bundles the required UPC++ runtime. | Monte Carlo Simulation, Liquid Crystal Membranes, Bilayer Simulation | Physical Sciences | Computational Software | Documentation, Uses, and more |
| microbecensus | mcc | 1.0 | MicroBeCensus, version 1.1.1, estimates the average genome size (AGS) of a microbial community directly from shotgun metagenomic reads and uses it to normalize gene abundances into biologically meaningful, cross-sample-comparable units. It works by rapidly mapping a subset of reads to a set of 30 universal single-copy marker genes; the coverage of these markers relative to total sequenced bases yields the average genome size, from which per-cell gene copy numbers can be computed. This normalization corrects for the fact that communities dominated by large genomes appear to have lower per-read functional-gene abundance, an artifact that AGS normalization removes. It is run on unassembled reads and produces the estimated AGS, genome equivalents, and the count of reads used, values that then feed downstream functional-abundance normalization. | Documentation, Uses, and more | |||
| microbecensus-sourceapp | mcc | 1.0 | This is a Python 3 packaging of MicrobeCensus (the 'SourceApp' variant, PyPI MicrobeCensus-SourceApp 1.1.2) used to estimate a metagenome's average genome size and to normalize gene abundances into genome-equivalent units. Like the original MicrobeCensus, it maps a subsample of shotgun reads to a panel of universal single-copy marker genes and infers average genome size from their relative coverage, then converts raw functional-gene counts into per-genome copy numbers that are comparable across samples of differing community composition. This SourceApp fork is maintained for compatibility as a dependency of the SourceApp microbial source-tracking pipeline and is packaged here to run under modern Python 3 environments. Usage mirrors the standard tool, taking shotgun reads and outputting the estimated average genome size and genome equivalents that drive downstream normalization. | Documentation, Uses, and more | |||
| migrate-n | lcc | N/A | Migrate-N is a population-genetics application that estimates effective population sizes and gene-flow (migration) rates between populations from genetic/sequence data using maximum-likelihood and Bayesian coalescent inference. Widely used in phylogeography and molecular ecology. | Documentation, Uses, and more | |||
| miniconda3 | mcc, lcc | N/A | Miniconda is a free minimal installer for conda. It is a small bootstrap version of Anaconda that includes only conda, Python, the packages they both depend on, and a small number of other useful packages (like pip, zlib, and a few others). If you need more packages, use the conda install command to install from thousands of packages available by default in Anaconda’s public repo, or from other channels, like conda-forge or bioconda. Description Source: https://docs.anaconda.com/free/miniconda/index.html |
Package Management, Dependency Management, Software Installation | Computer Science, Computer & Information Sciences | Package Manager | Documentation, Uses, and more |
| miniforge3 | mcc, lcc | N/A | Miniforge is a community-driven minimal installer for Conda, designed to provide a lightweight and flexible environment for managing packages and environments in Python and other languages. | Conda, Package Management, Environment Management, Python | Bioinformatics, Biological Sciences | Package Manager | Documentation, Uses, and more |
| minigraph | lcc | 1.0 | minigraph (from Heng Li's lh3/minigraph repository) is a sequence-to-graph mapper and incremental graph constructor used to build and align against pangenome reference graphs, with a focus on capturing large structural variation (typically 50 bp and above) across a collection of genomes. Starting from a linear reference, it incrementally augments the graph by aligning additional assemblies and adding their divergent segments as new nodes/edges, yielding a rGFA/GFA graph that encodes structural differences between haplotypes; it can likewise map long reads or contigs onto an existing graph, emitting alignments in the GAF format. It supports an incremental graph-construction mode from a reference plus additional assemblies as well as a long-read mapping mode against an existing graph. It is fast and memory-efficient, and is a common front end for pangenome pipelines (e.g., providing the SV backbone that tools like the Minigraph-Cactus workflow and vg build upon). It is oriented toward structural-variation-level pangenomics rather than base-level SNP graphs. | Documentation, Uses, and more | |||
| minimap | mcc | 1.0 | minimap2 is a fast, versatile pairwise sequence aligner that has become the standard mapper for long-read data while also handling short reads and full-length transcripts. It supports multiple presets tuned to different data types — map-ont and map-pb/map-hifi for noisy long genomic reads, splice for spliced mRNA/cDNA alignment (including long-read RNA-seq), asm5/asm10/asm20 for assembly-to-reference alignment, and sr for short reads — selected via a preset flag. It emits SAM/BAM alignments or, in a CIGAR-enabled mode, PAF-format pairwise mappings, and uses a minimizer-based seed-chain-align strategy that makes it both fast and memory-efficient on large genomes. This is version 2.22, a core component of nearly all long-read genomics, transcriptomics, and assembly pipelines. | alignment, genomics, bioinformatics, sequence analysis | Bioinformatics, Sequence Alignment | Command-line tool | Documentation, Uses, and more |
| minimap2 | mcc, lcc | 1.0 | minimap2 software. | Sequence Alignment, Bioinformatics, Genomics, Computational Biology | Bioinformatics, Biological Sciences | Bioinformatics Tool | Documentation, Uses, and more |
| miranda | mcc | 1.0 | miRanda (version 3.3a) predicts microRNA target sites by scanning miRNA sequences against 3'UTR or genomic sequences and scoring candidate binding sites on two complementary criteria: sequence complementarity (a position-weighted Smith-Waterman-style alignment that emphasizes the miRNA seed region) and the thermodynamic free energy of the predicted miRNA:mRNA duplex. Sites passing both a minimum alignment score and a maximum binding-energy threshold are reported as putative targets. Input is two FASTA files — the mature miRNA sequences and the target sequences — with tunable score and energy thresholds, producing per-site alignments, scores, energies, and coordinates. It is a long-established tool for computational miRNA target prediction, often used to generate candidate regulatory interactions for subsequent experimental validation or integration with expression data. | Documentation, Uses, and more | |||
| mirdeep | lcc | 1.0 | miRDeep2 (version 2.0.1.2 from bioconda) identifies known and novel microRNAs from small-RNA deep-sequencing data by modeling the characteristic hairpin (precursor) structure and the pattern of Dicer/Drosha processing that produces mature and star sequences. Its pipeline first maps collapsed small-RNA reads to a reference genome (via its mapper step using Bowtie), then scores candidate precursors using a probabilistic model of read-stacking and RNA secondary structure (RNAfold), reporting predicted miRNAs with confidence scores and read counts; a separate quantifier step quantifies expression against known miRBase mature/precursor sequences. Inputs include the small-RNA FASTQ/FASTA, a Bowtie-indexed genome, and optional miRBase reference sets for the species and related species. Outputs are HTML/CSV tables of novel and known miRNA predictions with structures and expression. It is a standard tool in small-RNA transcriptomics and miRNA discovery studies. | bioinformatics, microRNA, deep sequencing, RNA analysis | Molecular Biology, Biochemistry and Molecular Biology | Analysis Tool | Documentation, Uses, and more |
| mirdeep2 | mcc | 1.0 | miRDeep2, version 2.0.1.2, is a computational pipeline for discovering both known and novel microRNAs (miRNAs) from small-RNA deep-sequencing data by modeling the biogenesis of miRNAs from their precursor hairpins. It scores candidate loci using a probabilistic model of how Dicer processing distributes reads across a folded pre-miRNA hairpin (mature, star, and loop signatures), combined with the thermodynamic stability of the predicted secondary structure. The typical three-step workflow is a mapper stage that collapses and maps adapter-trimmed reads to the genome (producing an .arf alignment), a core identification-and-scoring stage that identifies and scores miRNAs against the genome and optional known-miRNA references (miRBase mature/precursor sets and related species), and a quantifier stage that measures expression of known miRNAs. Outputs include a ranked list of predicted miRNAs with confidence scores, structures, read stacks, and an HTML report. It remains one of the standard tools for miRNA annotation in both model and non-model organisms. | Bioinformatics, Computational Biology | Genetics, Biological Sciences | Tool | Documentation, Uses, and more |
| mitobim | lcc | 1.0 | MITObim (packaged from the chrishah/mitobim Docker image) is a tool for reconstructing complete or near-complete mitochondrial genomes directly from whole-genome next-generation sequencing data without needing a closely related, fully assembled mitogenome. It follows a baiting-and-iterative-mapping strategy built on top of the MIRA assembler: starting either from a related mitochondrial reference sequence or from a short seed (such as a barcoding gene), it iteratively recruits reads that overlap the growing assembly and extends the mitochondrial contig round by round until it stabilizes, effectively 'walking' out the mitogenome from the bulk sequencing data. Inputs are the sequencing reads (typically interleaved FASTQ) plus a seed/reference FASTA; outputs are the reconstructed mitochondrial contig(s) from each iteration. It is widely used in phylogenetics and biodiversity studies to obtain mitogenomes from genome-skimming or off-target reads. | bioinformatics, genomics, mitochondrial DNA, next-generation sequencing | Bioinformatics, Mitochondrial Genomics | Open-source | Documentation, Uses, and more |
| mitofinder | lcc | 1.0 | MitoFinder is a pipeline for assembling and annotating mitochondrial genomes (mitogenomes) from whole-genome or genome-skimming sequencing data, aimed at phylogenetics and molecular-ecology studies of non-model animals. It can run one of several assemblers (MEGAHIT, MetaSPAdes, or IDBA-UD) on the input reads, identify the mitochondrial contigs by similarity to a reference organelle genome, and then annotate protein-coding genes, tRNAs, and rRNAs, exporting results in formats such as GenBank, GFF, and FASTA ready for downstream alignment. It takes paired/single reads and a reference GenBank mitogenome, along with a specified genetic code. This is version 1.4. | Bioinformatics, Genomics, Mitochondrial Genomes, Next-Generation Sequencing, Annotation | Genomics, Biological Sciences | Annotation Tool | Documentation, Uses, and more |
| mitohifi | mcc, lcc | 1.0 | MitoHiFi (version 2.2, from the BioContainers image biocontainers/mitohifi:2.2_cv1) is a genome-assembly pipeline specialized in reconstructing complete, circular mitochondrial genomes from PacBio HiFi long reads or from an already-assembled set of contigs. Given a closely related reference mitogenome (which can be fetched automatically via its findMitoReference.py helper using an NCBI species query), it identifies candidate mitochondrial reads/contigs, assembles them with Hifiasm, and then circularizes and rotates the final sequence to start at a canonical gene. It filters out numts (nuclear mitochondrial insertions) and can annotate the resulting mitogenome using MitoFinder or MITOS. In reads mode it takes the HiFi reads together with the reference mitogenome sequence and annotation and produces a final FASTA plus annotation and coverage plots. It is widely used in reference genome projects (e.g., the Darwin Tree of Life) to recover organellar genomes as a byproduct of nuclear assembly. | Documentation, Uses, and more | |||
| mitoz | lcc | 1.0 | MitoZ is an all-in-one toolkit for animal mitochondrial genomics that performs de novo assembly, gene annotation, and visualization directly from raw sequencing reads. It can filter and quality-trim input reads, assemble the mitochondrial genome (with modes tuned for varying data quantity), annotate protein-coding genes, tRNAs, rRNAs, and the control region, and render a circular mitogenome map with per-base coverage. Provided here via a conda environment file (mitoz.yaml) rather than a channel package, it runs an all-in-one workflow over paired FASTQ reads with clade- and genetic-code-aware settings, emitting a GenBank-format annotation, a summary of recovered genes, and figures. Its clade- and genetic-code-aware annotation makes it popular for phylogenetics and DNA-barcoding studies across invertebrates and vertebrates. Outputs are directly usable for GenBank submission and comparative mitogenomic analysis. | bioinformatics, genomics, mitochondrial analysis | Bioinformatics, Genomics | Analysis Tool | Documentation, Uses, and more |
| mkfontdir | mcc | N/A | mkfontdir is a utility that generates an index of font files in a directory, creating a fonts.scale file that is used by the X Window System to manage fonts. | Font Management, X Window System, Utilities | Software Engineering, Other Computer and Information Sciences | Utility | Documentation, Uses, and more |
| mkfontscale | mcc | N/A | mkfontscale is a utility that generates an index of scalable fonts for the X Window System. | Font Management, X Window System, Utilities | Software Engineering, Other Computer and Information Sciences | Utility | Documentation, Uses, and more |
| mkl | mcc, lcc, ecc | N/A | Intel Compiler Family (C/C++/Fortran for x86_64) | Mathematical Library, Optimized Computations, Thread-Safe Functions | Computer Science | Computational Software | Documentation, Uses, and more |
| mkl32 | mcc, lcc | N/A | The Intel® oneAPI Math Kernel Library (oneMKL) helps you achieve maximum performance with a math computing library of highly optimized, extensively parallelized routines for CPU and GPU. The library has C and Fortran interfaces for most routines on CPU, and SYCL interfaces for some routines on both CPU and GPU. Description Source: https://www.intel.com/content/www/us/en/developer/tools/oneapi/onemkl.html#gs.4ofwv10 |
Mathematics, Computational Software, Hpc Tools | Other Mathematics | Library | Documentation, Uses, and more |
| mmg | mcc, lcc | N/A | MMG is a software package designed for the generation and adaptation of unstructured meshes in two and three dimensions. It is particularly useful for numerical simulations in scientific computing. | mesh generation, mesh adaptation, scientific computing, numerical simulations | Numerical Methods, Computational Science | Open Source | Documentation, Uses, and more |
| mmseqs2 | lcc | 1.0 | MMseqs2 (Many-against-Many sequence searching) is an ultra-fast and sensitive open-source suite for searching and clustering very large protein and nucleotide sequence sets, achieving BLAST-level sensitivity at orders-of-magnitude greater speed through a prefiltering k-mer matching stage followed by vectorized local alignment. Its capabilities include homology search, all-against-all clustering at chosen sequence-identity thresholds (including the linear-time linclust algorithm), taxonomy assignment, and profile searches; the companion Linclust scales to billions of sequences. Data is organized as MMseqs2 databases created from input sequences, and results are converted back to BLAST-tabular or FASTA with its conversion utilities. It underlies tools such as ColabFold's MSA generation. This is version 13.45111. | Sequence Searching, Sequence Clustering, Bioinformatics, Computational Biology | Biological Sciences | Sequence Search & Clustering Tool | Documentation, Uses, and more |
| moddotplot | mcc | 1.0 | ModDotPlot draws dot plots of sequence self identity and pairwise identity at whole genome scale, and is designed for looking at the internal structure of highly repetitive DNA such as centromeric satellite arrays. Rather than comparing every base, it reduces each interval of the sequence to a sketch of selected k-mers called modimizers, which approximates average nucleotide identity between all pairs of intervals quickly enough to plot an entire chromosome on a laptop scale of memory. It runs in a static mode that writes plots and an identity matrix for use in batch pipelines, and an interactive mode that serves a browser application for panning and zooming through a plot at changing resolution. Inputs are FASTA files, optionally with BED annotation tracks to draw alongside the plot, and a previously computed identity matrix can be replotted without recomputing it. Outputs are the self identity matrix as a paired end BED file plus triangle, full matrix and histogram plots in both raster and vector form, which makes it useful both for exploring repeat organisation and for producing publication figures. | Documentation, Uses, and more | |||
| modelaveraging-deeplearning | lcc | 1.0 | This is a curated GPU deep-learning research environment rather than a single program, built on the NVIDIA NGC PyTorch 24.04 base (PyTorch 2.3, CUDA 12.4) and assembled for work on conditional generative models and pretrained-model averaging/ensembling. It combines HuggingFace transformers (4.44.2) and diffusers (0.30.3) for transformer and diffusion models, a set of normalizing-flow libraries (normflows, nflows, zuko) for density estimation and generative modeling, image backbones (timm, open_clip_torch), and a suite of generative-model evaluation metrics—clean-fid, pytorch-fid, torch-fidelity for FID/IS-style image-quality scores, plus geomloss and POT for optimal-transport/Wasserstein distances. NumPy is pinned to 1.24.4 for compatibility across these packages. It targets researchers training, sampling from, and quantitatively evaluating generative models on GPU, providing a consistent, conflict-free stack from a single environment. | Documentation, Uses, and more | |||
| modelaveraging-r | lcc | 1.0 | This is a curated R environment (built on rocker/geospatial R 4.5.1) assembled to reproduce and support covariate-dependent stacking / ensemble spatial-prediction analyses, notably the CovariateDependentStacking study (arXiv:2408.09755). Rather than a single package, it bundles a coherent toolchain for model averaging and ensemble prediction: Bayesian estimation (MCMCpack), flexible regression and smoothing (gam, np for nonparametric methods, mgcv-style GAMs), machine-learning base learners (kernlab, randomForest, and the spatially aware RandomForestsGLS), spatial statistics (spdep, spatialreg, GWmodel for geographically weighted models), visualization (ggplot2, viridis, tidyr), timing (tictoc), and the GitHub-only gamFactory package (mfasiolo/gamFactory) plus remotes to install it. The workflow it enables trains multiple base predictive models and combines them with weights that themselves depend on covariates/location ('stacking'), yielding improved, spatially adaptive predictions. It is aimed at statisticians and spatial-data researchers doing ensemble prediction, and is packaged so the study's code runs with its exact dependency set. | Documentation, Uses, and more | |||
| molpro | lcc | N/A | Molpro software. | Quantum Chemistry, Computational Chemistry, High-Performance Computing | Chemical Sciences, Physical Chemistry | Commercial | Documentation, Uses, and more |
| momap | mcc, lcc | N/A | momap software. | Documentation, Uses, and more | |||
| mono | lcc | N/A | mono software. | cross-platform, open-source, development, framework | Computer Science, Software Engineering | Framework | Documentation, Uses, and more |
| moose | lcc | 1.0 | MOOSE (Multiphysics Object-Oriented Simulation Environment) is Idaho National Laboratory's open-source, massively parallel finite-element framework for solving tightly coupled systems of nonlinear partial differential equations. Built on the libMesh and PETSc libraries, it uses the Jacobian-Free Newton-Krylov (JFNK) method and provides a modular, object-oriented architecture in which physics are expressed as "Kernels," "BoundaryConditions," "Materials," and other pluggable objects, letting developers assemble and couple multiphysics simulations (heat conduction, solid mechanics, phase field, reactive transport, Navier-Stokes, etc.) via its extensive physics modules. Simulations are defined in a human-readable HIT/GetPot input file and run in parallel via MPI, with mesh adaptivity, automatic differentiation, and restart support. This build compiles MOOSE from source against INL's moose-dev conda toolchain (2023.11.30). It is used to build custom engineering and scientific simulation applications, particularly in nuclear engineering and materials, and typically requires compiling a derived application against the framework. | Simulation, Multiphysics, Finite Element Method | Engineering & Technology | Computational Software | Documentation, Uses, and more |
| mosdepth | mcc | 1.0 | mosdepth (version 0.3.8 from bioconda) is a fast tool for computing sequencing depth and coverage from BAM or CRAM alignment files, written to be substantially quicker than samtools/bedtools-based approaches by scanning each region only once. It can report mean coverage over fixed-size windows, over user-supplied BED regions (e.g., exome targets or genes), or per base, and it produces distribution summaries used to assess coverage uniformity and callable fractions. Useful capabilities include windowed or regional depth reporting, a fast mode that skips per-base detail for speed, quantized and threshold outputs for callable-region reporting, and MAPQ/flag filtering, yielding compact bgzipped BED and global/region distribution text files. It is a standard QC step in whole-genome, exome, and targeted resequencing pipelines and is a favorite for large cohort coverage summaries. | genomics, bioinformatics, sequencing, coverage | Bioinformatics, Biological Sciences | Command Line Tool | Documentation, Uses, and more |
| mothur | mcc, lcc | 1.0 | mothur (version 1.48.0) is a single, self-contained software package for processing and analyzing microbial community amplicon sequencing data such as 16S rRNA and ITS. It offers an integrated command set covering the full pipeline: sequence quality screening and trimming, alignment against reference databases (e.g. SILVA), chimera detection (VSEARCH/UCHIME), OTU clustering and/or phylotype/ASV-style analysis, taxonomic classification, and a broad range of downstream ecology statistics including alpha- and beta-diversity metrics, rarefaction, ordination (PCoA/NMDS), and hypothesis tests (AMOVA, HOMOVA, metastats, LEfSe). It runs interactively at its own command prompt or in batch mode from a script of commands, and historically implements the widely cited Schloss lab SOPs. Inputs are sequence reads and reference alignment/taxonomy files; outputs include shared/OTU tables and diversity summaries. It remains a popular alternative to QIIME 2 for reproducible microbiome analysis in a single tool. | Bioinformatics, Microbial Communities, High-Throughput Sequencing | Microbial Ecology, Biological Sciences | Analysis Tool | Documentation, Uses, and more |
| mpc | mcc | N/A | Gnu Mpc is a C library for the arithmetic of complex numbers with arbitrarily high precision and correct rounding of the result. | Cryptographic Protocol, Secure Computation, Privacy-Preserving, Collaborative Computation | Computer & Information Sciences | Cryptographic Protocol | Documentation, Uses, and more |
| mpdaf-3.5-fitsio | lcc | 1.0 | MPDAF, the MUSE Python Data Analysis Framework, is an astronomy library for working with data produced by the VLT/MUSE integral-field spectrograph, which yields 3D data cubes carrying a full optical spectrum at every spatial pixel. It provides Python objects for the fundamental data types (Cube, Image, and Spectrum) with methods for reading/writing FITS files, extracting sub-cubes and apertures, collapsing along spatial or spectral axes, computing narrow-band images, fitting and measuring emission/absorption lines, performing continuum and sky handling, managing world-coordinate and wavelength calibration, and building and manipulating source catalogs. This environment is MPDAF version 3.5 from PyPI paired with the conda fitsio 1.1.7 package, which supplies fast, direct FITS I/O via CFITSIO for reading and writing the large FITS cubes and tables MUSE analyses involve. It is provided within a multi-tool Miniconda base container that bundles several bioinformatics and astronomy packages, each in its own conda/pip environment; typical use is importing mpdaf in Python to load and analyze a MUSE cube. | astronomy, data analysis, FITS, spectroscopy, MUSE, Python, scientific computing, astrophysics | "", Astronomy and Planetary Sciences | Library | Documentation, Uses, and more |
| mpfr | mcc, lcc, ecc | N/A | The MPFR library is a C library for multiple-precision floating-point computations with correct rounding. MPFR has continuously been supported by the INRIA and the current main authors come from the Caramba and AriC project-teams at Loria (Nancy, France) and LIP (Lyon, France) respectively; see more on the credit page. MPFR is based on the GMP multiple-precision library. Description Source: https://www.mpfr.org/ |
Arbitrary-Precision Arithmetic, Floating-Point Computation, Mathematical Library | Numerical Analysis, Mathematics | Computational Software | Documentation, Uses, and more |
| mpi | mcc, lcc, ecc | N/A | Intel(R) MPI Library | Hpc, Parallel Computing, Message Passing, Scalable Computing | Computer Science, Computer & Information Sciences | Library | Documentation, Uses, and more |
| mpi4py | lcc | N/A | Mpi4py provides Python bindings for the Message Passing Interface (MPI) standard. It is implemented on top of the MPI specification and exposes an API which grounds on the standard MPI-2 C++ bindings. Description Source: https://pypi.org/project/mpi4py/ |
Parallel Computing, Message Passing Interface, Python Library | Computer Science, Computer & Information Sciences, Other Computer & Information Sciences | Python Library | Documentation, Uses, and more |
| mpich | mcc, lcc, ecc | N/A | MPICH MPI implementation (device=ch3:nemesis) | Mpi Implementation, Parallel Computing, High Performance Computing | Computer Science, Computer & Information Sciences | Library | Documentation, Uses, and more |
| mpip | lcc | N/A | mpiP: a lightweight profiling library for MPI applications. | Software Utility, Hpc, Mpi Libraries, Cluster Management, Installation Tool | Engineering & Technology | Documentation, Uses, and more | |
| mrbayes | mcc, lcc | 1.0 | MrBayes (bioconda 3.2.7) is a program for Bayesian inference of phylogeny and evolutionary-model parameters, estimating the posterior distribution over tree topologies, branch lengths, and substitution-model parameters using Metropolis-coupled Markov Chain Monte Carlo (MC^3). It supports a wide range of substitution models for DNA, RNA, protein, restriction-site, and morphological (standard) data, including partitioned analyses that apply different models to different data subsets, relaxed molecular clocks, and divergence-time estimation. Analyses are specified in a NEXUS file containing the aligned data and a mrbayes command block that sets the model (lset/prset), MCMC parameters (ngen, nchains, samplefreq), and output options; the program produces sampled trees (.t) and parameter samples (.p) and summarizes them via its sump and sumt commands (a consensus tree with posterior-probability support and convergence diagnostics such as the average standard deviation of split frequencies). It can be run in serial or MPI-parallel builds and is one of the most widely cited tools for Bayesian phylogenetics. | Bayesian Inference, Phylogenetics, Evolutionary Biology | Systematics & Population Biology, Biological Sciences | Software Tool | Documentation, Uses, and more |
| msmc | mcc | 1.0 | Documentation, Uses, and more | ||||
| msmc-decode | mcc | 1.0 | Documentation, Uses, and more | ||||
| msmc-tools | mcc | 1.0 | Documentation, Uses, and more | ||||
| msmc2 | mcc | 1.0 | Documentation, Uses, and more | ||||
| multiqc | mcc, lcc | 1.0 | MultiQC is a reporting tool that scans a directory tree of analysis outputs and aggregates the quality-control results and logs from many different bioinformatics tools across all samples into a single, self-contained interactive HTML report. It recognizes output formats from well over a hundred tools — FastQC, fastp, Cutadapt, STAR, HISAT2, Bowtie, Salmon, samtools, Picard, GATK, bcftools, Qualimap, featureCounts, Kraken, and many more — parsing their logs automatically and building comparative plots and summary tables that let a user spot outlier samples at a glance. It is run over a top-level results directory, with options to name the report, point at specific directories, force overwrite, or restrict/ignore particular modules. The output is an interactive report (with plots that can be exported) plus a multiqc_data folder of parsed numbers for programmatic reuse. It is the de facto standard final QC-summary step in sequencing pipelines. This is version 1.28. | Bioinformatics, Computational Biology, Data Analysis, Reporting Tool | Biological Sciences | Visualization Tool | Documentation, Uses, and more |
| mumax3 | lcc | 1.0 | mumax3 is a GPU-accelerated micromagnetics simulator used in condensed-matter physics and spintronics research to model magnetization dynamics in nano- and micro-scale ferromagnetic structures. It numerically solves the Landau-Lifshitz-Gilbert equation on a finite-difference cuboid mesh, accounting for exchange, demagnetizing (magnetostatic), anisotropy, Zeeman, and spin-transfer/spin-orbit torque terms. Simulations are described in a compact scripting language defining geometry, material regions, field/current excitations, and solver settings, and are launched by running the simulator on a script file; outputs include time-series tables of averaged quantities and OVF field snapshots that can be visualized or converted. This build is the CUDA 12.0 Linux binary (v3.11.1) and runs on NVIDIA GPUs, with JupyterLab available for interactive post-processing of the results. | micromagnetics, GPU computing, simulation, magnetic materials | Condensed Matter Physics | Simulation Software | Documentation, Uses, and more |
| mummer4 | lcc | 1.0 | MUMmer 4 is a fast, versatile suite for whole-genome and large-scale sequence alignment based on efficient suffix-array/maximal-exact-match anchoring, suitable for aligning entire chromosomes and multi-megabase genomes. Its core drivers are nucmer (nucleotide-vs-nucleotide alignment) and promer (translated six-frame protein-level alignment), which anchor on maximal exact matches and then extend/cluster them; downstream utilities such as delta-filter, show-coords, show-snps, show-aligns, and mummerplot filter alignments and report coordinates, SNPs/indels, and generate dotplots. Typical use runs nucmer to align a query against a reference, then reports coordinates with show-coords and produces a dotplot with mummerplot. This particular environment bundles MUMmer4 (v4.0.1) together with a broader comparative-genomics alignment toolkit: minimap2 2.30 (long-read/genome mapping), CrossMap 0.7.3 (coordinate liftover between assemblies), paf2chain 0.1.1 (convert PAF alignments to UCSC chain format), Mugsy 1.2.3 (multiple whole-genome alignment), LAST 1642 (adaptive-seed alignment), and the legacy MUMmer 3.23, giving users a single place for pairwise and multiple genome alignment and coordinate conversion. | Bioinformatics, Sequence Alignment, Genomics | Genomics, Biological Sciences | Alignment Tool | Documentation, Uses, and more |
| mumps | lcc | N/A | A MUltifrontal Massively Parallel Sparse direct Solver | Linear Algebra, Sparse Matrices, Computational Science, Engineering | Applied Mathematics, Mathematics | Solver | Documentation, Uses, and more |
| munge | mcc | N/A | Documentation, Uses, and more | ||||
| muscle | mcc, lcc | 1.0 | MUSCLE (MUltiple Sequence Comparison by Log-Expectation), version 3.8.1551, is a fast and widely cited program for constructing multiple sequence alignments of proteins or nucleotides. This 3.8 release uses an iterative refinement strategy: it builds a rapid initial progressive alignment from k-mer distance estimates, refines the guide tree from the first alignment, and then performs tree-dependent restricted partitioning to improve the alignment, balancing accuracy against speed with options to limit iterations for large datasets. Input and output are FASTA (with support for other alignment formats). MUSCLE is a common alternative to MAFFT and Clustal for producing alignments that feed phylogenetic inference, profile construction, and conservation analysis. (Note that this is the classic v3.8 command-line interface, which differs from the redesigned MUSCLE5.) | Sequence Alignment, Bioinformatics | Biological Sciences | Bioinformatics Tool | Documentation, Uses, and more |
| mvapich2 | mcc, lcc | N/A | OSU MVAPICH2 MPI implementation | Mpi Implementation, High Performance Computing, Infiniband, Open-Source | Computer Science, Physical Sciences | Library/Tool | Documentation, Uses, and more |
| mvp | mcc | 1.0 | MVP (Metagenomic Viral Pipeline, distributed as the conda package 'mvip') is an automated pipeline for recovering and characterizing viruses from metagenomic assemblies. It chains together established virus-identification and quality tools to detect viral contigs, assess their completeness and quality, dereplicate/cluster them into viral operational taxonomic units (vOTUs), and summarize their abundance and taxonomy across samples. It takes assembled contigs (and read sets for coverage estimation) as input and produces curated viral genome sets with per-sample tables and reports suitable for viral community ecology. The pipeline is run through its mvip command-line modules that execute the successive detection, quality-control, and clustering stages. This is version 1.1.5. | Documentation, Uses, and more | |||
| namd | lcc | 1.0 | NAMD is a high-performance parallel molecular-dynamics engine developed at the University of Illinois for simulating large biomolecular systems—proteins, membranes, nucleic acids, and their solvent environments—often into the millions of atoms. Built on the Charm++ parallel runtime, it scales efficiently across many nodes and GPUs and is the standard MD backend paired with the VMD visualization tool and the CHARMM/AMBER force fields. This deployment provides the vendor-prebuilt 3.0.2 verbs-SMP-CUDA builds, meaning it uses native InfiniBand (verbs) for multi-node communication combined with per-node multicore threading and CUDA GPU offload for maximum throughput. Input is a configuration file referencing a PSF topology, coordinate (PDB), and force-field parameter files; a typical cluster run is launched through charmrun or the verbs launcher with the desired per-node process and GPU-device settings. It outputs trajectory (DCD), restart, and energy logs for equilibration, free-energy, and steered-MD studies. | Molecular Dynamics, Biomolecular Systems, Parallel Computing, Cuda, Simulation | Biochemistry and Molecular Biology | Computational Software | Documentation, Uses, and more |
| nanocomp | mcc | 1.0 | NanoComp, part of the NanoPack suite for Oxford Nanopore (and PacBio) long-read data, generates side-by-side comparisons of multiple sequencing runs, samples, or barcodes to evaluate and contrast their quality and yield. From several input datasets it computes summary statistics (read count, total bases, N50, mean/median read length and quality) and renders comparative visualizations such as violin/box plots of read-length and quality distributions, cumulative-yield-over-time curves, and log-transformed length plots, all collected into a single HTML report and a stats table. It accepts multiple input types, including FASTQ files, aligned BAM files, or Nanopore sequencing_summary.txt files, with per-dataset names supplied for labeling. Version 1.25.6 is provided within the nanopack conda environment of a Rocky 9 long-read/genomics container focused on Nanopore/PacBio QC, classification, and assembly polishing. | Documentation, Uses, and more | |||
| nanofilt | mcc | 1.0 | NanoFilt is a lightweight command-line tool from the NanoPack suite for filtering and trimming Oxford Nanopore long-read sequencing data. It reads FASTQ (optionally gzipped) from standard input and applies user-specified thresholds to discard low-quality or short reads and to trim bases from read ends, which is a common preprocessing step before assembly or mapping of Nanopore data. Key options set a minimum average read quality, minimum and maximum read length, and a fixed number of bases to crop from the start or end of each read; it can also use the sequencing_summary file for quality information. Because it operates as a Unix filter, it is typically used within a decompress-to-recompress pipe over a FASTQ stream. Note that NanoFilt is deprecated upstream in favor of chopper, but remains widely used; this is bioconda nanofilt 2.8.0 in the nanopack environment. | Bioinformatics, Sequence Analysis, Nanopore Sequencing | Bioinformatics, Biological Sciences | Tool | Documentation, Uses, and more |
| nanoplot | mcc | 1.0 | NanoPlot is a plotting and summary-statistics tool for long-read sequencing data from Oxford Nanopore and PacBio platforms, part of the NanoPack suite. It ingests FASTQ files, aligned BAM/CRAM files, or Nanopore sequencing_summary.txt files and produces a set of publication-quality plots — read-length histograms, cumulative yield, read-length-versus-quality bivariate plots, and (from BAMs) alignment identity distributions — together with a text/HTML report of key metrics such as N50, mean/median read length and quality, and total yield. Version 1.46.2 is a standard first-look QC step for long-read runs, letting researchers quickly judge run quality, throughput, and read-length characteristics before assembly or mapping. | Bioinformatics, Genomics, Sequencing Data, Quality Control, Visualization | Bioinformatics, Biological Sciences | Bioinformatics Tool | Documentation, Uses, and more |
| nanostat | mcc | 1.0 | NanoStat is a lightweight command-line utility, part of the NanoPack toolkit for long-read sequencing, that computes and reports summary statistics describing an Oxford Nanopore or PacBio dataset. It reports metrics such as the total number of reads and bases, mean and median read length, read-length N50, mean and median read quality, and the number of reads/bases above various length and quality thresholds. It can read several input types — FASTQ files, aligned BAM, or sequencing_summary.txt files from the Nanopore basecaller — making it easy to profile a run either from raw reads or directly from the basecaller output, printing a compact table to standard output. It is commonly used as a quick QC snapshot before and after read filtering in long-read pipelines. This is version 1.6.0, installed within a nanopack conda environment. | Documentation, Uses, and more | |||
| nasm | mcc, lcc | N/A | Netwide Assembler (NASM) is an asssembler for the x86 CPU architecture portable to nearly every modern platform with code generation for many platforms old and new. Description Source: https://www.nasm.us/ |
Assembler, Cross-Platform, X86, Programming | Computer Science, Computer & Information Sciences | Assembler | Documentation, Uses, and more |
| ncbi-rmblastn | mcc | N/A | RMBlast is a tool for aligning nucleotide sequences against a database of reference sequences, optimized for use with the NCBI BLAST algorithm. | bioinformatics, sequence alignment, NCBI, BLAST | Bioinformatics, Biological Sciences | Command-line tool | Documentation, Uses, and more |
| ncl | lcc | 1.0 | NCL (the NCAR Command Language) is an interpreted, domain-specific programming language and toolkit developed at NCAR for the analysis and visualization of scientific data, used extensively in the atmospheric, oceanographic, and climate-science communities. It natively reads the self-describing data formats common in geosciences — NetCDF (3 and 4), HDF4/HDF5, GRIB1/GRIB2, and shapefiles — and provides hundreds of built-in functions for gridded-data manipulation, regridding and interpolation, spatial and temporal averaging, spectral and EOF/statistical analysis, and climate-diagnostic computations. Its integrated graphics engine produces publication-quality contour maps, vector/streamline plots, cross-sections, and XY plots with extensive map-projection support. Scripts are written in the NCL language and executed with the interpreter, or run interactively. Although NCL is now in maintenance mode with the community migrating toward the Python GeoCAT ecosystem, it remains widely used for existing climate-diagnostics workflows. This build is version 6.6.2 from conda-forge. | Scientific Data Analysis, Data Visualization, Atmospheric Sciences | Environmental Sciences, Earth & Environmental Sciences | Scripting Language | Documentation, Uses, and more |
| ncurses | mcc, lcc, ecc | N/A | The ncurses (new curses) library is a free software emulation of curses in System V Release 4.0 (SVr4), and more. It uses terminfo format, supports pads and color and multiple highlights and forms characters and function-key mapping, and has all the other SVr4-curses enhancements over BSD curses. SVr4 curses became the basis of X/Open Curses. Description Source: https://invisible-island.net/ncurses/announce.html |
Programming Library, Text-Based User Interfaces, Terminal Applications | Computer Science, Computer & Information Sciences | Programming/Development | Documentation, Uses, and more |
| ncview | lcc | 1.0 | Ncview is a lightweight visual browser for netCDF files, giving scientists a quick, interactive look at the contents of gridded multidimensional datasets common in climate, atmospheric, oceanographic, and other geophysical modeling. It opens a GUI (X11) listing the variables in the file; selecting one displays a color-mapped 2D field with a movie-style control panel to animate through the time (or other) dimension, change color maps and ranges, and click any grid cell to read its value or plot a 1D time series/line through the data. It is meant for rapid data inspection and quality checking rather than producing publication figures, letting users confirm that model output or observational data looks sensible before deeper analysis. Version 2.1.7 depends on the netCDF and X11 libraries and is a staple utility in Earth-system-science workflows. | Data Visualization, Netcdf, Earth System Data | Earth and Environmental Sciences | Data Analysis Tool | Documentation, Uses, and more |
| netcdf | lcc | N/A | C Libraries for the Unidata network Common Data Form | Scientific Data, Data Storage, Data Access, Data Formats, Data Manipulation, Data Visualization | Environmental Data, Other Natural Sciences | Library | Documentation, Uses, and more |
| netcdf-c | mcc | N/A | The Unidata network Common Data Form (netCDF) is an interface for scientific data access and a freely-distributed software library that provides an implementation of the interface. The netCDF library also defines a machine-independent format for representing scientific data. Together, the interface, library, and format support the creation, access, and sharing of scientific data. Description Source: https://docs.unidata.ucar.edu/netcdf-c/current/ |
Data Format, Scientific Data, Data Access, Data Sharing | Climate and Global Dynamics, Other Natural Sciences | Data Format | Documentation, Uses, and more |
| netcdf-cxx | lcc | N/A | C++ Libraries for the Unidata network Common Data Form | Netcdf, C++ | Earth & Environmental Sciences | Data Processing | Documentation, Uses, and more |
| netcdf-fortran | mcc, lcc | N/A | Fortran Libraries for the Unidata network Common Data Form | Scientific Data, Fortran Interface, Netcdf Library | Data Analysis, Other Natural Sciences | Programming Interface | Documentation, Uses, and more |
| netcdf-intel | lcc | N/A | NetCDF-Intel is a high-performance implementation of the NetCDF data format, optimized for Intel architectures. It provides tools for reading and writing array-oriented scientific data in a platform-independent manner. | Data Format, Scientific Computing, High Performance Computing, Data Management | Data Analysis, Computer and Information Sciences | Library | Documentation, Uses, and more |
| netlib-scalapack | mcc, lcc | N/A | ScaLAPACK is a library of high-performance linear algebra routines for parallel distributed memory machines. ScaLAPACK solves dense and banded linear systems, least squares problems, eigenvalue problems, and singular value problems. Description Source: https://www.netlib.org/scalapack/ |
linear algebra, parallel computing, high-performance computing, distributed systems | Numerical Analysis, Applied Mathematics | Library | Documentation, Uses, and more |
| netlogo | mcc | N/A | NetLogo software. | Modeling, Simulation, Agent-based, Education | Social Sciences, Other social sciences | Simulation Software | Documentation, Uses, and more |
| networkx | lcc | N/A | NetworkX is a Python package for the creation, manipulation, and study of the structure, dynamics, and functions of complex networks. Description Source: https://networkx.org/ |
Python Library, Graph Theory, Network Analysis, Data Visualization | Computer & Information Sciences | Computational Software | Documentation, Uses, and more |
| nextflow | mcc, lcc | 1.0 | Nextflow is a workflow-management system and domain-specific language for composing complex, data-intensive computational pipelines that are portable and reproducible across environments. Built on a dataflow programming model, it connects processes through asynchronous channels so that steps execute in parallel as soon as their inputs are ready, and it separates pipeline logic from execution so the same workflow runs unchanged on a laptop, an HPC scheduler (SLURM, PBS, LSF), Kubernetes, or the cloud, with software provisioned via conda, Docker, or Singularity. It provides built-in resumability (caching completed tasks so a run can be resumed), execution reports and timelines, and seamless integration with the nf-core community pipeline collection. Version 25.04.4 is a mainstay for reproducible bioinformatics and general scientific data analysis at scale. | Workflow Manager, Computational Biology, Data Science | Computational Biology, Biological Sciences | Open Source | Documentation, Uses, and more |
| nextpolish | mcc | 1.0 | NextPolish is a genome-assembly polishing tool that corrects base-level errors (mismatches and small insertions/deletions) in draft assemblies produced from long reads. It can polish using short reads (Illumina) alone, long reads (Nanopore/PacBio) alone, or a hybrid combination, iterating alignment and consensus scoring over multiple rounds to improve per-base accuracy without altering large-scale structure. It is configured through a run configuration file listing the input assembly, the read datasets (via a file-of-file-names), and the sequence of polishing tasks/rounds, and it emits the polished FASTA. It is commonly applied after long-read assembly and before or alongside scaffolding. This is version 1.4.1. | Bioinformatics, Computational Biology, Long-Read Sequencing, Error Correction | Genetics, Biological Sciences | Sequence Analysis Tool | Documentation, Uses, and more |
| nextpolish2 | mcc | 1.0 | NextPolish2 is a genome-assembly polishing tool designed to correct residual base-level errors in long-read (particularly PacBio HiFi) assemblies while, critically, preserving true heterozygosity rather than erasing it. It uses accurate short or HiFi k-mer datasets (built with yak) as a truth reference and is repeat-aware, so it avoids over-correcting repetitive regions—addressing weaknesses of earlier consensus polishers. The bundled minimap2 aligns reads to the draft assembly, samtools handles the alignment files, and yak constructs the k-mer databases used for correction. A typical workflow maps HiFi reads to the draft with minimap2 and then polishes the assembly using the resulting alignment together with the k-mer databases. Version 0.2.2 is used at the final polishing stage of high-quality (near telomere-to-telomere) diploid genome assemblies. | Documentation, Uses, and more | |||
| ngenomesyn | mcc | 1.0 | NGenomeSyn (N Genome Synteny) is a visualization tool for depicting synteny and structural relationships across multiple genomes in a single publication-quality figure. Given pairwise alignment or synteny/link files (for example from minimap2, MUMmer, or a structural-variant caller like SyRI) together with the genomes' sequence lengths, it draws stacked genome bars connected by colored ribbons that show collinear blocks, inversions, translocations, and other rearrangements, with extensive control over layout, spacing, colors, and the addition of feature tracks (gene density, GC content, etc.). It is driven by a plain-text configuration file listing the input links and display options and rendered to an output figure. It is widely used in comparative and structural genomics to communicate multi-genome collinearity and chromosomal rearrangements clearly, especially for more than two genomes where circular plots become cluttered. | Documentation, Uses, and more | |||
| nghttp2 | ecc | N/A | This is an implementation of the Hypertext Transfer Protocol version 2 in C. The framing layer of HTTP/2 is implemented as a reusable C library. On top of that, we have implemented an HTTP/2 client, server and proxy. We have also developed load test and benchmarking tools for HTTP/2. An HPACK encoder and decoder are available as a public API. | Http/2, Web Protocol, Networking | Computer Science, Computer & Information Sciences | Library | Documentation, Uses, and more |
| nginx | lcc | 1.0 | nginx (version 1.25.3 from conda-forge) is a high-performance, event-driven HTTP server, reverse proxy, and load balancer that handles large numbers of concurrent connections with a small memory footprint. In an HPC/container context it is typically used to serve static content, proxy requests to backend application servers, terminate TLS, and cache responses. Its behavior is governed by a declarative nginx.conf organized into http, server, and location blocks, and it supports virtual hosting, gzip compression, URL rewriting, rate limiting, and upstream health-aware balancing. It runs as a master process with worker processes and can be reloaded gracefully without dropping connections. Here it is bundled within a multi-application conda container, exposed as its own app entry point alongside independent bioinformatics environments. | Web Server, Reverse Proxy, Load Balancer, Open Source | Computer Science | Web Server | Documentation, Uses, and more |
| ngsdist | lcc | 1.0 | ngsDist estimates pairwise genetic distances between individuals directly from genotype likelihoods rather than from called genotypes, making it well suited to low- and medium-coverage next-generation sequencing data where hard genotype calls are uncertain. By propagating genotype uncertainty (typically taken from ANGSD-produced genotype likelihoods in beagle format), it produces less biased distance estimates that can then feed distance-based phylogenetic reconstruction (e.g., FastME) or clustering/ordination. It supports bootstrapping over sites/blocks to assess support for the resulting tree topology. Inputs are a genotype-likelihood file plus the number of individuals and sites (and optional labels); the output is a pairwise distance matrix. It is part of the fgvieira population-genomics toolset and is installed here from its GitHub repository. | Documentation, Uses, and more | |||
| ngsld | mcc | 1.0 | ngsLD estimates pairwise linkage disequilibrium (LD) between genetic sites while explicitly accounting for genotype uncertainty, making it suited to low- and medium-coverage population-genomics data analyzed in the ANGSD/genotype-likelihood framework. Instead of hard-called genotypes it takes per-site genotype likelihoods (e.g. beagle-format output from ANGSD) plus a sites/positions file and computes LD statistics such as r^2, D, and D' for pairs of sites within a specified maximum distance. It ships helper scripts (fit_LDdecay.R, prune_graph) for fitting LD-decay curves and for LD-based pruning of markers to obtain approximately independent SNP sets. It is built from source and run as a command-line C++ program; this container also provides an R environment with LDheatmap for visualizing the resulting LD structure. | Ngs, Bioinformatics, Genomics, Sequencing | Genetics, Biological Sciences | Bioinformatics Tool | Documentation, Uses, and more |
| ngstools | lcc | 1.0 | ngsTools is a suite of programs by Matteo Fumagalli and collaborators for population-genetic analysis of low- and medium-coverage next-generation sequencing data that works directly from genotype likelihoods rather than hard-called genotypes, thereby propagating sequencing uncertainty into the analyses. The recursive clone bundles several components: ngsPopGen (estimates of nucleotide diversity, FST, PCA, and site-frequency spectra), ngsF and ngsF-HMM (per-individual inbreeding coefficients), ngsLD (linkage-disequilibrium estimation from genotype likelihoods), and ngsDist (pairwise genetic distances for tree/PCA building). It integrates tightly with ANGSD, which typically produces the genotype-likelihood (.glf/.beagle) and allele-frequency inputs that ngsTools then consumes; the container also ships samtools/bcftools/htslib and supporting R packages for downstream visualization. The tools are individually invoked, making the suite a mainstay of conservation and evolutionary genomics on non-model organisms sequenced at low depth. | Documentation, Uses, and more | |||
| ngsutils | lcc | 1.0 | NGSUtils is a suite of command-line tools for manipulating and analyzing next-generation sequencing data across the common BAM, BED, FASTQ, GTF, and VCF file formats. It is organized into several toolsets invoked as subcommands: bamutils (BAM filtering, tagging, coverage, RNA-seq counting, statistics), fastqutils (FASTQ QC, trimming, filtering, format checks), bedutils (interval manipulation, sorting, overlap), and gtfutils, among others. These utilities cover many small but frequently needed genomics operations — for instance BAM filtering, read counting against annotations, FASTQ name extraction, or BED interval extension — that glue together larger analysis pipelines. Inputs and outputs are the corresponding standard sequencing file formats, with commands following a toolset-then-command pattern. This is bioconda ngsutils 0.5.9, a Python 2-era but still-useful general-purpose NGS Swiss-army knife. | Ngs, Bioinformatics, Sequencing, Data Analysis | Bioinformatics, Biological Sciences | Research Tools | Documentation, Uses, and more |
| ninja | mcc, lcc, ecc | 1.0 | NINJA, in this container, is the large-scale neighbor-joining phylogenetic inference tool bundled inside the Dfam TE Tools image at /opt/NINJA/Ninja (note: it is the phylogenetics program, not the Ninja build system). It implements an exact but heavily optimized, disk-and-memory-efficient neighbor-joining algorithm that can build trees from distance matrices or alignments containing tens of thousands to hundreds of thousands of sequences, far beyond what naive NJ implementations handle. Within the TE Tools ecosystem it is used by RepeatModeler to cluster and build guide trees for large sets of repeat-family sequences during transposable-element discovery. It is run as a Java-backed executable that takes an alignment or a precomputed distance matrix and emits a Newick tree, with options controlling the input type and clustering method. It is included primarily as a dependency of the Dfam repeat-annotation workflow but can be used standalone for very large NJ tree building. | Build System, Software Development, Automation | Software Engineering, Systems & Development, Engineering & Technology | Utility | Documentation, Uses, and more |
| nnlojet | lcc | 1.0 | Documentation, Uses, and more | ||||
| nnlojet-run | lcc | 1.0 | Documentation, Uses, and more | ||||
| node | lcc | N/A | Node.js JavaScript runtime (LTS, glibc-2.17 build) | Javascript, Server-Side Scripting, Cross-Platform | Computer Science, Computer & Information Sciences | Programming Platform | Documentation, Uses, and more |
| node-js | mcc, ecc | N/A | Node.js is a JavaScript runtime built on Chrome's V8 JavaScript engine, enabling developers to build scalable network applications. | JavaScript, Runtime, Server-side, Event-driven, Network applications | Software Engineering, Computer Science | Open Source | Documentation, Uses, and more |
| nonpareil | mcc | 1.0 | Nonpareil (version 3.5.5 from bioconda) estimates the coverage and sequencing effort of metagenomic datasets by quantifying read redundancy, allowing researchers to judge how completely a microbial community has been sampled and how much additional sequencing would be needed to reach a target coverage. It works by measuring the fraction of reads that have overlapping/similar partners within a dataset, using either the original k-mer-based algorithm or a faster alignment-free approach, and fitting a redundancy-versus-effort curve. The command-line tool produces a redundancy (.npo) file that is then loaded into the companion Nonpareil R package to compute the Nonpareil diversity index (Nd) and project the sequencing effort required for near-complete coverage. It is widely used to compare community complexity across samples and to plan the depth of metagenomic surveys. Outputs support both per-sample coverage estimates and cross-sample diversity comparisons. | metagenomics, assembly analysis, bioinformatics | Bioinformatics, Biological Sciences | Analysis Tool | Documentation, Uses, and more |
| notebook | mcc, lcc | 1.0 | This is a Jupyter Notebook environment (classic Notebook 6.4.8) providing an interactive, browser-based interface for writing and running Python code in cells alongside rich text, equations, and inline plots, well suited to exploratory data analysis and reproducible reporting on the cluster. It is bundled with the core scientific-Python data stack: pandas for tabular data manipulation, NumPy for array computing, scikit-learn for classical machine learning (classification, regression, clustering, model selection), and matplotlib for plotting, plus IPython/ipykernel for the interactive kernel. Users launch the notebook server (typically forwarding a port from the compute node) and work in .ipynb notebooks. It provides a general-purpose analysis sandbox rather than a domain-specific tool. | Documentation, Uses, and more | |||
| notung | lcc | N/A | Notung is a phylogenetics application for reconciling gene trees with species trees. It performs duplication/loss inference, rooting, rearrangement of weakly supported branches, and dating of gene family evolution events. Widely used in comparative genomics and molecular evolution studies to reconcile discordant gene and species trees. Version 2.9 is provided as a Java-based tool. | Phylogenetics, Evolutionary Biology, Bioinformatics | Computational Biology, Bioinformatics | Application | Documentation, Uses, and more |
| novoplasty | lcc | 1.0 | NOVOPlasty, version 4.2, is a de novo organelle-genome assembler and heteroplasmy detector that reconstructs mitochondrial and chloroplast genomes directly from whole-genome sequencing reads using a seed-and-extend strategy. Starting from a single seed sequence — a related organelle region, a conserved gene, or even a portion of the reads — it iteratively extends and circularizes the organelle genome by finding overlapping reads, without assembling the far larger nuclear genome. It is configured entirely through a plain-text config file specifying genome type/range, k-mer size, the seed and (optionally) reference sequences, and the paired-end read files, then run against that config. Outputs include the circular (or contiguous) organelle assembly, a log, and, when applicable, detected heteroplasmic variants. It is especially popular for plant plastomes and animal mitogenomes where a closely related seed is available. | bioinformatics, genomics, assembly, next-generation sequencing | Bioinformatics, Genomics | Assembly Tool | Documentation, Uses, and more |
| npm | mcc, ecc | N/A | Documentation, Uses, and more | ||||
| nri-md | lcc | 1.0 | NRI-MD applies Neural Relational Inference (NRI), an unsupervised graph neural network approach, to molecular-dynamics (MD) trajectories in order to learn the latent interaction structure among protein residues. Rather than relying on simple distance or correlation cutoffs, the NRI encoder-decoder framework infers a dynamic interaction graph while simultaneously learning to predict future residue positions, so the recovered edges reflect functionally relevant couplings that drive conformational motion. Inputs are typically time series of residue coordinates (for example Cα positions extracted from an MD trajectory), and outputs include the inferred adjacency/interaction network and associated edge-type probabilities that can be visualized as an allosteric or communication network. Built on PyTorch and deep-learning tooling, it is intended for researchers studying allostery, signal propagation, and dynamic residue-residue communication in proteins. The repository provides training and evaluation scripts for the encoder/decoder models along with utilities to convert trajectories into the expected input format. | Documentation, Uses, and more | |||
| ntlink | mcc | 1.0 | ntLink is a lightweight genome-assembly scaffolding tool that orders, orients, and joins draft contigs using long reads (Nanopore or PacBio) together with a memory-efficient minimizer-based mapping approach rather than full alignment. It computes minimizer sketches of the long reads and the assembly, finds shared minimizers to infer which contigs are linked and in what relative orientation and distance, and builds a scaffold graph to produce longer scaffolds, with an optional iterative rounds and gap-filling capability that can also patch sequence into the gaps between joined contigs. It is notably fast and low-memory compared to alignment-based scaffolders. Inputs are a draft assembly FASTA and long-read FASTA/FASTQ; output is a scaffolded assembly plus the mapping/graph intermediates, produced through its Snakemake-driven scaffold interface. Version 1.3.10 is provided as a conda environment in a Rocky 8 Miniconda container of metagenomics/assembly tools that also includes a Space Ranger binary. | Documentation, Uses, and more | |||
| numactl | mcc, lcc | N/A | numactl is simple NUMA policy support. It consists of a numactl program to run other programs with a specific NUMA policy and a libnuma shared library ("NUMA API") to set NUMA policy in applications. Description Source: https://github.com/numactl/numactl |
Numa, Memory Management, Performance Optimization | Other Computer & Information Sciences, Engineering & Technology | Tool | Documentation, Uses, and more |
| numba | lcc | N/A | Numba is an open source JIT compiler that translates a subset of Python and NumPy code into fast machine code. Description Source: https://numba.pydata.org/ |
Jit Compiler, Python Optimization, Numpy Optimization | Computer Science, Computer & Information Sciences | Open-Source | Documentation, Uses, and more |
| numpy | lcc | N/A | NumPy is the fundamental package for scientific computing in Python. It is a Python library that provides a multidimensional array object, various derived objects (such as masked arrays and matrices), and an assortment of routines for fast operations on arrays, including mathematical, logical, shape manipulation, sorting, selecting, I/O, discrete Fourier transforms, basic linear algebra, basic statistical operations, random simulation and much more. Description Source: https://numpy.org/doc/stable/user/whatisnumpy.html |
Scientific Computing, Numerical Computing, Data Analysis, Mathematics | Mathematics, Computer & Information Sciences | Computational Software | Documentation, Uses, and more |
| nvhpc | lcc | N/A | The NVIDIA HPC Software Development Kit (SDK) includes the proven compilers, libraries and software tools essential to maximizing developer productivity and the performance and portability of HPC applications. Description Source: https://developer.nvidia.com/hpc-sdk |
Hpc, Gpu Acceleration, Compilers, Libraries, Development Tools | Computer Science, Computer & Information Sciences | Development Tool | Documentation, Uses, and more |
| nvidia-rapids | lcc | 1.0 | NVIDIA RAPIDS is an open-source suite of GPU-accelerated libraries that mirror the familiar PyData APIs so that data-science and machine-learning workloads run on NVIDIA GPUs with minimal code changes. The 24.12 release bundled here includes cuDF (a pandas-compatible GPU DataFrame), cuML (scikit-learn-style machine learning), cuGraph (graph analytics), and related components, alongside CuPy (a NumPy-like GPU array library), a CUDA-enabled PyTorch build, and JupyterLab for interactive work. Running on CUDA 12.5, it lets users accelerate ETL, feature engineering, clustering, regression/classification, and graph algorithms by orders of magnitude over CPU equivalents, often by simply swapping a pandas import for its cuDF equivalent. It is intended for end-to-end accelerated analytics pipelines on GPU nodes, and the included JupyterLab makes it convenient for exploratory, notebook-based development. | GPU Computing,Data Science,Machine Learning,Open Source | Data Analytics, Computer Science | Library | Documentation, Uses, and more |
| nwchem | mcc, lcc | 1.0 | NWChem is a comprehensive, massively parallel computational chemistry package developed at PNNL for modeling molecular and periodic systems on hardware ranging from workstations to supercomputers. It covers a broad range of methods: Gaussian-basis quantum chemistry (Hartree-Fock, DFT, MP2, coupled-cluster up to CCSD(T), multireference approaches, TDDFT for excited states), plane-wave density functional theory and ab initio molecular dynamics through its NWPW module, and classical/QM-MM molecular dynamics. Its parallelism is built on the Global Arrays toolkit over MPI, giving good scaling of both computation and distributed memory for large calculations. Users supply a text input deck specifying the geometry, basis set, charge/multiplicity, theory level, and task (energy, optimize, frequencies, dynamics), typically run under MPI across multiple processes. This deployment wraps the official NWChem developer ORAS image (nwchem-dev built against OpenMPI 4.1.x) so the MPI-enabled binary runs inside a container. | Computational Chemistry, High-Performance Computing, Open-Source Software | Chemical Sciences | Computational Chemistry Software | Documentation, Uses, and more |
| ocelot | lcc | N/A | Documentation, Uses, and more | ||||
| ocelot_api | lcc | N/A | Documentation, Uses, and more | ||||
| oclfpga | mcc, lcc | N/A | The Intel® oneAPI DPC++/C++ Compiler provides optimizations that help your applications run faster on Intel® 64 architectures on Windows* and Linux*, with support for the latest C, C++, and SYCL language standards. This compiler produces optimized code that can run significantly faster by taking advantage of the ever-increasing core count and vector register width in Intel® Xeon® processors and compatible processors. Description Source: https://www.intel.com/content/www/us/en/docs/dpcpp-cpp-compiler/get-started-guide/2024-0/overview.html |
Opencl, Fpga, Acceleration | Engineering & Technology | Framework | Documentation, Uses, and more |
| ocr | lcc | N/A | Open Community Runtime (OCR) for shared memory | Documentation, Uses, and more | |||
| octopus | lcc | 1.0 | Octopus (version 16.3) is a scientific electronic-structure code specializing in real-space, real-time (time-dependent) density-functional theory. It represents wavefunctions on a real-space grid rather than a basis set and is designed to compute the ground state and, especially, the time-dependent response and dynamics of electrons in molecules, clusters, nanostructures, and periodic solids under external fields. Capabilities include TDDFT for optical absorption and excited states, real-time propagation of the Kohn-Sham equations, calculation of polarizabilities, hyperpolarizabilities, and other response properties, and simulation of light-matter interaction and non-linear phenomena. It interfaces with the libxc exchange-correlation library and supports hybrid MPI+OpenMP parallelism for HPC use; this build is compiled from source on CentOS 7.6 with a full numerical stack (FFT, BLAS/LAPACK, etc.). Input is a single text inp file specifying the system, grid, and calculation mode, and it is run from the working directory. It is widely used in computational condensed-matter and materials physics. | Quantum Mechanics, Condensed Matter Physics, Simulation Software | Physics, Physical Sciences | Simulation | Documentation, Uses, and more |
| ohpc | mcc, lcc | N/A | OpenHPC is a collaborative, community-driven effort to provide a comprehensive and cohesive open-source HPC stack to the community. Through close collaboration with key stakeholders across the industry, OpenHPC aims to provide a reference collection of open-source HPC software components and best practices to enable the development, testing, and deployment of HPC systems. Ultimately, building on the resecarch outputs from various related fields. | Hpc, High Performance Computing, Open-Source, Collaborative, Community-Driven | Computer Science, Engineering & Technology | Computational Software | Documentation, Uses, and more |
| ollama | lcc, ecc | 1.0 | Ollama is a lightweight runtime for downloading, managing, and serving large language models on local hardware. It provides a server with an OpenAI-compatible HTTP API and a model library covering open-weight families such as Llama, Qwen, DeepSeek, Gemma, and Mistral, handling model storage, quantized formats, and GPU acceleration automatically. On ECC the container detects and uses the node NVIDIA GPUs, and the shared model directory is preconfigured so commonly used models are not re-downloaded per user. Typical use is interactive chat or programmatic inference against the served API from Python or curl, including embedding generation for retrieval workflows. Input is a prompt or API request; output is generated text or embeddings returned to the caller. | AI, Machine Learning, Natural Language Processing, Local Deployment | Command Line Tool | Documentation, Uses, and more | |
| omb | mcc, lcc | N/A | OSU Micro-benchmarks | Documentation, Uses, and more | |||
| open-webui | ecc | 1.0 | Open WebUI is a self-hosted, extensible web front-end for interacting with large language models entirely under local/institutional control. It speaks both the OpenAI-compatible API and the Ollama API, so it can drive locally running open models or remote OpenAI-style endpoints, and it provides a polished chat interface with conversation history, multiple model switching, prompt templates, and role-based multi-user access. A key feature is offline document RAG (retrieval-augmented generation): users can upload files or point to a document collection and have the model answer questions grounded in that content without data leaving the environment. In this deployment it is delivered as a per-user OOD Batch Connect kiosk app on ECC — launched in a self-contained noVNC/Firefox stack — giving researchers a private in-browser AI Chat on cluster resources. It is installed from the open-webui PyPI package. | Documentation, Uses, and more | |||
| openbabel | mcc, lcc | 1.0 | Open Babel is a chemical toolbox and software library designed to speak the many languages of chemical data, enabling conversion, searching, analysis, and manipulation of molecular structures across more than 100 file formats used in computational chemistry, cheminformatics, and molecular modeling. Its command-line workhorse obabel converts between formats (e.g. SMILES, InChI, MOL/SDF, PDB, XYZ, CIF), generates 2D and 3D coordinates, adds hydrogens, performs energy minimization with built-in force fields, computes molecular fingerprints and descriptors, and filters or substructure-searches sets of molecules — for instance reading SMILES input and producing 3D SDF structures. It also provides a Python API (pybel) for scripting these operations. It is a foundational utility for preparing and interconverting chemical data in docking, QSAR, and simulation workflows. This is conda-forge openbabel 3.1.1. | Cheminformatics, Molecular Modeling, File Conversion, Open Source | Cheminformatics, Molecular Modeling | Library | Documentation, Uses, and more |
| openbabel311 | mcc | 1.0 | Open Babel (version 3.1.1, built from source from the openbabel GitHub repository into /usr/local/openbabel-3-1-1) is a chemical toolbox for interconverting, searching, filtering, and analyzing molecular data across a very large number of chemical file formats (over a hundred), including SMILES, InChI, MOL/SDF, PDB, MOL2, XYZ, and CIF. Its primary command-line interface, obabel, converts between formats, generates 3D coordinates and adds hydrogens, performs energy minimization with built-in force fields, computes molecular descriptors and fingerprints, and enables substructure searching and conformer generation. It also exposes Python and C++ APIs for programmatic cheminformatics. It is a workhorse utility in computational chemistry and drug-discovery pipelines for preparing and standardizing molecular inputs for docking, QM, and virtual-screening tools. This build compiles the library from source on an Ubuntu 22.04 conda container. | Chemistry, File Conversion, Molecular Modeling, Open Source | Chemical Sciences, Chemical Informatics | Library | Documentation, Uses, and more |
| openblas | mcc, lcc | N/A | An optimized BLAS library based on GotoBLAS2 | Linear Algebra, Blas Library, Optimized Functions, Open Source | Mathematics, Computer & Information Sciences | Library | Documentation, Uses, and more |
| opencoarrays | lcc | N/A | ABI to leverage the parallel programming features of the Fortran 2018 DIS | Fortran, Parallel Computing, Coarrays, High Performance Computing | Computer Science | Library | Documentation, Uses, and more |
| openfoam | mcc, lcc | 1.0 | OpenFOAM is the free, open source CFD software developed primarily by OpenCFD Ltd since 2004. It has an extensive range of features to solve anything from complex fluid flows involving chemical reactions, turbulence and heat transfer, to acoustics, solid mechanics and electromagnetics. Description Source: https://www.openfoam.com/ |
Computational Fluid Dynamics, Open-Source Software, Finite Volume Method, Hpc | Fluid Dynamics, Physical Sciences | Computational Software | Documentation, Uses, and more |
| openjdk | mcc, ecc | N/A | OpenJDK is a free, Open-Source version of the Java Development Kit for the Java Platform, Standard Edition (Java SE). Description Source: https://openjdk.org/ |
Open-Source, Java Development, Software Development, Programming | Software Engineering, Computer & Information Sciences | Compiler | Documentation, Uses, and more |
| openmm | mcc, lcc | 1.0 | OpenMM is a high-performance toolkit and library for molecular dynamics (MD) simulation, designed around a flexible Python API backed by highly optimized CPU and GPU (CUDA/OpenCL) compute platforms; this build pins conda-forge OpenMM 8.1.0. It lets users define custom force fields and arbitrary custom forces, run explicit- and implicit-solvent simulations, apply thermostats/barostats, and perform enhanced-sampling and free-energy calculations, either as a standalone engine or as a library embedded in larger workflows. Because forces and integrators can be expressed programmatically, it is popular both for production biomolecular MD and for methods development. This particular environment additionally bundles OpenMiChroM, which builds on OpenMM to run chromatin/chromosome polymer models (e.g., the Minimal Chromatin Model) for simulating 3D genome organization. A typical use is a short Python script that loads a PDB and force field, creates a System and Simulation object, minimizes energy, and integrates dynamics while reporting to DCD/PDB and state-data logs. | Molecular Dynamics, Biomolecular Simulations, Computational Chemistry, Biophysics, Gpu Acceleration | Bioinformatics, Biological Sciences | Molecular Dynamics Software | Documentation, Uses, and more |
| openmpi | mcc, lcc | N/A | The Open MPI Project is an open source implementation of the Message Passing Interface (MPI) specification that is developed and maintained by a consortium of academic, research, and industry partners. Open MPI is therefore able to combine the expertise, technologies, and resources from all across the High Performance Computing community in order to build the best MPI library available. Description Source: https://www.open-mpi.org/ |
Mpi Library, Parallel Computing, High Performance Computing, Distributed Memory System | Computer Science, Computer & Information Sciences | Development Tools | Documentation, Uses, and more |
| openmpi3 | lcc | N/A | A powerful implementation of MPI | MPI, Parallel Computing, High Performance Computing, Cluster Computing | Computer Science | Library | Documentation, Uses, and more |
| openmpi4 | mcc, lcc | N/A | A powerful implementation of MPI/SHMEM | HPC, Parallel Computing, Message Passing, Distributed Systems | Computer Science | Library | Documentation, Uses, and more |
| opensees | lcc | N/A | OpenSees (Open System for Earthquake Engineering Simulation, v2.5.0) is a finite-element framework for simulating the response of structural and geotechnical systems to earthquakes and other loads. It supports nonlinear static and dynamic analysis and is a standard research tool in earthquake and structural engineering. Provided here as a Singularity container. | Simulation, Finite Element Analysis, Earthquake Engineering, Structural Analysis, Geotechnical Engineering | Civil Engineering, Engineering & Technology | Finite Element Analysis (Fea) | Documentation, Uses, and more |
| openssh | mcc, lcc | N/A | OpenSSH is the premier connectivity tool for remote login with the SSH protocol. It encrypts all traffic to eliminate eavesdropping, connection hijacking, and other attacks. In addition, OpenSSH provides a large suite of secure tunneling capabilities, several authentication methods, and sophisticated configuration options. Description Source: https://www.openssh.com/ |
Ssh, Secure Networking, Utilities | Computer Science, Computer & Information Sciences | Utility | Documentation, Uses, and more |
| openssl | mcc, lcc, ecc | N/A | OpenSSL software is a robust, commercial-grade, full-featured toolkit for general-purpose cryptography and secure communication. Description Source: https://www.openssl.org/ |
Cryptography, Security, Encryption, Networking | Cryptography, Computer & Information Sciences, Other Computer & Information Sciences | Toolkit | Documentation, Uses, and more |
| opera-ms | mcc | N/A | OPERA-MS is a hybrid metagenome assembler for reconstructing microbial genomes from mixed communities. It combines short-read (Illumina) and long-read (Oxford Nanopore / PacBio) data to scaffold and assemble strain-resolved genomes from complex metagenomic samples. Used in microbiome and environmental genomics research for producing near-complete draft genomes and resolving repeats that short reads alone cannot span. Installed here as a Conda environment layered on Miniconda3. | Documentation, Uses, and more | |||
| operams | mcc, lcc | 1.0 | OPERA-MS is a hybrid metagenome assembly and scaffolding pipeline that combines the complementary strengths of accurate short reads and long-range long reads to reconstruct near-complete genomes from complex microbial communities. It first uses a short-read (or provided) assembly, then leverages long reads (Nanopore/PacBio) to scaffold contigs across repeats and to cluster and disentangle contigs belonging to different species and strains, addressing the multi-genome, variable-abundance nature of metagenomes. It bundles a full supporting stack, including MUMmer 3.23 for alignment, kraken2 for taxonomic binning/species clustering, SPAdes 3.13.0 and MEGAHIT for the underlying assembly, plus R, and integrates CheckM to assess the completeness and contamination of the resulting genome bins. It takes paired short reads together with a long-read set and produces reconstructed genome bins in an output directory. It is deployed as an Ubuntu 18.04 container carrying OPERA-MS and its full bundled toolchain. | Documentation, Uses, and more | |||
| optislang | lcc | N/A | Ansys optiSLang (23R2) is a process-integration and design-optimization application for engineering simulation. It provides robust design optimization, sensitivity analysis, design of experiments, parametric studies, and uncertainty/reliability quantification, and it orchestrates workflows across Ansys and third-party solvers. Engineers use it to automate multidisciplinary optimization and to improve product robustness and reliability. | Optimization, Engineering, Simulation, Sensitivity Analysis | Mechanical Engineering, Optimization | Commercial | Documentation, Uses, and more |
| orca | mcc, lcc | N/A | orca software. | Quantum Chemistry, Computational Chemistry, Quantum Mechanics | Physical Chemistry, Chemical Sciences | Quantum Chemistry Software | Documentation, Uses, and more |
| orthofinder | mcc, lcc | 1.0 | OrthoFinder is a comparative-genomics platform that, given the complete proteomes of a set of species, infers the full set of orthogroups (the genes descended from a single gene in the last common ancestor of the species), and from these resolves orthologs and paralogs, gene trees, a rooted species tree, and gene-duplication events mapped onto that tree. It runs an integrated pipeline of all-versus-all sequence search (DIAMOND by default, or BLAST), normalized clustering into orthogroups (the original OrthoFinder algorithm), gene-tree inference, species-tree estimation (STAG/STRIDE for rooting), and duplication/orthology assignment, and it produces comparative statistics and orthogroup-level summaries. Input is simply a directory of one protein FASTA per species, and it yields orthogroup tables, gene and species trees, and orthologue lists. Version 3.1.0 is provided as its own conda environment/app in a Rocky 9 Miniconda base container that also bundles single-cell, viral/microbial genomics, and population-genetics tools. | Bioinformatics, Phylogenomics, Orthology, Gene Set Inference | Bioinformatics, Biological Sciences | Genome Analysis Software | Documentation, Uses, and more |
| orthologer | mcc | 1.0 | Orthologer is the orthology-delineation software that powers the OrthoDB database, developed by the EZlab, provided here from the upstream ezlabgva/orthologer v3.9.0 Docker image (repackaged as an MCC Singularity container with a fix to a return-code bug in common_odb.sh). It identifies orthologous groups of genes across species by combining best-reciprocal-hit clustering with graph-based grouping, and supports two principal modes: ODB-mapper, which maps a user's protein set onto the precomputed OrthoDB hierarchy to assign genes to known orthologous groups (useful for annotation transfer and completeness assessment), and de novo orthology inference within a user-supplied collection of FASTA proteomes. It is driven through the orthologer command and helper scripts that set up a project directory, run all-vs-all comparisons, and cluster the results, emitting orthologous-group assignments and per-species membership tables. It is commonly used to define single-copy ortholog sets for phylogenomics and to benchmark gene-set completeness. | Documentation, Uses, and more | |||
| os | mcc, ecc | N/A | The os module in Python provides a way of using operating system dependent functionality like reading or writing to the file system. | Python, Operating System, File Management | Software Engineering, Computer Science | Library | Documentation, Uses, and more |
| p4vasp | lcc | N/A | p4vasp | Documentation, Uses, and more | |||
| p7zip | ecc | N/A | p7zip is a port of 7za.exe for POSIX systems like Unix (Linux, Solaris, OpenBSD, FreeBSD, Cygwin, AIX, ...), MacOS X and also for BeOS and Amiga. Description Source: https://p7zip.sourceforge.net/ |
File Archiver, Compression Software, Command-Line Tool, Open-Source Software | Computer Science, Computer & Information Sciences | Command-line Tool | Documentation, Uses, and more |
| paml | lcc | 1.0 | PAML (Phylogenetic Analysis by Maximum Likelihood), version 4.9, is Ziheng Yang's package for maximum-likelihood and Bayesian analysis of DNA and protein sequences on a fixed phylogeny, best known for detecting signatures of natural selection. Its central programs are codeml, which fits codon-substitution models to estimate the nonsynonymous/synonymous rate ratio (dN/dS, ω) and run site, branch, and branch-site tests for positive selection, and baseml for nucleotide-level models; the package also includes evolver (sequence simulation) and mcmctree (Bayesian divergence-time estimation with fossil calibrations). Analyses are configured through a control file (e.g. codeml.ctl) specifying the alignment, tree, model, and null/alternative hypotheses; likelihood-ratio tests between nested models then assess significance. PAML is a standard tool in molecular evolution for studying adaptive protein evolution and estimating divergence times, though it requires careful setup and pre-aligned, in-frame sequences. | Phylogenetics, Evolutionary Analysis, Maximum Likelihood, Phylogenetic Inference | Genetics, Biological Sciences | Computational Software | Documentation, Uses, and more |
| pango | mcc | N/A | pango is a library for layout and rendering of text, with an emphasis on internationalization. Pango can be used anywhere that text layout is needed; however, most of the work on Pango so far has been done using the GTK widget toolkit as a test platform. Pango forms the core of text and font handling for GTK. Description Source: https://github.com/GNOME/pango |
Text Layout, Font Handling, Internationalization | Computer & Information Sciences | Text Processing | Documentation, Uses, and more |
| panpatch | mcc | 1.0 | panpatch is a tool by Glenn Hickey that improves the contiguity of fragmented genome assemblies by leveraging a pangenome graph built from multiple assemblies of the same sample together with a reference. By aligning several draft assemblies (for example different haplotype-resolved or independently produced assemblies of one individual) into a graph, panpatch identifies where one assembly can supply sequence that fills gaps or joins broken contigs in another, effectively 'patching' the fragments toward more contiguous, potentially telomere-to-telomere (T2T) sequences. It outputs a TSV report describing the patching decisions, an optional BED of the patched intervals, and patched per-haplotype FASTA files representing the improved assembly. It operates on pangenome-graph inputs (such as GFA) and is aimed at assembly-curation workflows where multiple sequence sources are available for the same genome. This is the checksum-verified static binary panpatch v0.3 running on Rocky 9. | Documentation, Uses, and more | |||
| papi | mcc, lcc | N/A | Performance Application Programming Interface | Performance Monitoring, Api, Hardware Counter | Computer Science | Library | Documentation, Uses, and more |
| parallel | mcc, lcc | 1.0 | GNU parallel is a shell tool for executing jobs concurrently, letting users saturate all cores of a machine (or distribute work across several machines over SSH) instead of running commands one at a time in a serial loop. It reads a list of inputs—from arguments, a file, or standard input—and runs a command template once per input, substituting each item with {} (with variants like {.} to strip an extension or {/} for the basename), while controlling the number of simultaneous jobs with -j. It preserves and orders output, can build combinations of multiple input lists, supports dry-run, retries, job logs, and resuming, and is a drop-in accelerator for embarrassingly parallel tasks such as converting many files or running an analysis over many samples. This is version 20220722; it is a ubiquitous utility for maximizing throughput in HPC and data-processing shell workflows. | Parallel Computing, Command-Line Tool, Task Scheduling | Computer & Information Sciences | Command-Line Tool | Documentation, Uses, and more |
| parallel-forkmanager | mcc | 1.0 | Parallel::ForkManager is a Perl module that provides a simple way to run work in parallel by forking child processes, and is a very common dependency of bioinformatics pipelines written in Perl. The programmer sets a maximum number of simultaneous children and then wraps a loop body between a start and a finish call; the module takes care of forking, of blocking once the limit is reached, and of reaping children as they complete, so a script processing many files or many genomic regions can use a whole node without hand written process management. Callbacks can be registered to run when each child starts or finishes, and a child can pass a data structure back to the parent as it exits, which makes it straightforward to collect per task results without temporary files. It is a library rather than a command line program, so it is used by running your own Perl script with the interpreter that provides it, which in this container is simply perl. | Documentation, Uses, and more | |||
| parallel-netcdf | mcc | N/A | PnetCDF (Parallel netCDF) is a high-performance parallel I/O library for accessing files in format compatibility with Unidata's NetCDF, specifically the formats of CDF-1, 2, and 5. Description Source: https://parallel-netcdf.github.io/ |
High-Performance I/O, Big Data, Parallel Computing | Earth & Environmental Sciences, Physical Sciences | I/O Library | Documentation, Uses, and more |
| paraview | mcc, lcc | 1.0 | ParaView 5.10.1 OSMesa build for server-side CPU/software rendering on any MCC compute node. | Visualization, Data Analysis, Open-Source, Multi-Platform | Data Visualization, Physical Sciences, Engineering & Technology, Natural Sciences, Computer & Information Sciences, Other Natural Sciences | Data Visualization | Documentation, Uses, and more |
| parmetis | mcc | N/A | ParMETIS is an MPI-based parallel library that implements a variety of algorithms for partitioning unstructured graphs, meshes, and for computing fill-reducing orderings of sparse matrices. ParMETIS extends the functionality provided by METIS and includes routines that are especially suited for parallel AMR computations and large scale numerical simulations. Description Source: http://glaros.dtc.umn.edu/gkhome/metis/parmetis/overview |
Graph Partitioning, Sparse Matrix Ordering, Scientific Computing, Parallel Computing | Computer Science, Computer & Information Sciences | Computational Tool | Documentation, Uses, and more |
| partitionfinder | lcc | 1.0 | PartitionFinder selects the best-fit data-partitioning scheme and corresponding nucleotide or amino-acid substitution models for phylogenetic analyses of multi-locus sequence alignments. Given an alignment and a set of user-defined data blocks (e.g. genes, or 1st/2nd/3rd codon positions), it searches over ways to group those blocks into partitions and, for each candidate scheme, evaluates substitution models, ranking schemes by AIC, AICc, or BIC to balance fit against overparameterization. This produces partitioning and model choices that can be handed directly to downstream tree-inference programs such as RAxML, IQ-TREE, MrBayes, or BEAST. It is configured through a partition_finder.cfg file (specifying the alignment, data blocks, branch-length linkage, model set, and search algorithm — e.g. greedy or rcluster). This build runs under a Python 2.7 conda environment as required by the v2.1.1 codebase. | phylogenetics, model selection, bioinformatics | Computational Biology, Bioinformatics | Desktop Application | Documentation, Uses, and more |
| patch | ecc | N/A | The `patch` utility is a tool used to apply changes to files based on a diff file, which contains the differences between two versions of a file or set of files. | version control, file management, diff, software development | Software Engineering, Other Computer and Information Sciences | Command-line tool | Documentation, Uses, and more |
| patchelf | mcc, lcc | N/A | PatchELF is a small utility to modify the dynamic linker and RPATH of ELF executables. | Dependency Management, Executable Modification, Dynamic Linker Paths | Systems and Development, Computer & Information Sciences | Library Management | Documentation, Uses, and more |
| pathofact | lcc | N/A | PathoFact is a bioinformatics pipeline for the prediction of virulence factors, bacterial toxins, and antimicrobial resistance genes in metagenomic and genomic datasets. It combines multiple prediction tools (including signal peptide detection via SignalP) into a unified workflow, and is commonly used in microbiome and pathogen genomics research to characterize the pathogenic potential of microbial communities. | Documentation, Uses, and more | |||
| pb-assembly | lcc | 1.0 | pb-assembly is PacBio's official FALCON-based toolkit for de novo genome assembly and diploid phasing from PacBio long reads. It wraps the FALCON assembler (for generating a haploid primary contig assembly via overlap detection and consensus), FALCON-Unzip (to phase the assembly into primary contigs plus associated haplotigs representing the alternate allele), and supporting utilities. It is driven by configuration files that specify input read FASTA/FOFN and stage-specific parameters, and it orchestrates the compute-intensive daligner overlap, error-correction, string-graph, and consensus steps, typically on distributed/cluster resources. The output is a primary contig set with phased haplotigs, useful for producing partially phased diploid assemblies of eukaryotic genomes; it was designed for noisy PacBio continuous-long-read (CLR) subreads. | Documentation, Uses, and more | |||
| pbccs | mcc, lcc | 1.0 | pbccs (the ccs tool) generates PacBio HiFi reads by computing highly accurate circular-consensus sequences from the multiple subread passes of a single SMRTbell molecule. Because a HiFi read is the consensus of many passes around the same circularized template, ccs delivers long reads (typically 10-25 kb) with QV20+ (99%+) accuracy, combining long-read length with short-read-like base quality. Input is a subreads (or unaligned) PacBio BAM; output is a BAM of consensus reads carrying per-read quality and pass-count tags, filterable by minimum predicted accuracy and number of passes. Large runs are commonly chunked and merged for parallel processing. Its HiFi output is the foundation of modern high-accuracy long-read genome assembly, variant calling, and Iso-Seq workflows. | Documentation, Uses, and more | |||
| pbjelly | mcc, lcc | 1.0 | PBJelly, a component of the PBSuite, is a genome upgrading and gap-closing tool that uses PacBio long reads to fill captured gaps and extend contigs in existing draft assemblies. It runs a staged pipeline — setup, mapping (via the BLASR long-read aligner), support, extraction, assembly of gap-spanning reads, and output — to identify reads that span or flank scaffold gaps and then locally assemble them to close or shrink those gaps. Configuration is driven by a Protocol.xml file specifying the reference, reads, and BLASR parameters, and the pipeline is invoked stage by stage. It is a legacy Python 2.7 tool; in this build it runs in a Python 2.7 environment with bioconda BLASR 5.3.2 and networkx 1.9. This is PBSuite PBJelly version 15.8.24. | bioinformatics, genomics, DNA assembly | Bioinformatics, Genomics | Assembly Tool | Documentation, Uses, and more |
| pbmm2 | mcc, lcc | 1.0 | pbmm2 is PacBio's official SMRT wrapper around the minimap2 aligner, tuned to align native PacBio data (subreads, CCS/HiFi reads, and Iso-Seq transcripts) while preserving PacBio-specific BAM tags and read-group metadata. Version 1.13.1. It accepts PacBio BAM, FASTA, FASTQ, or a dataset XML as input and writes sorted, indexed BAM output directly, eliminating the separate minimap2-to-samtools sort step. Alignment behavior is selected through preset modes — SUBREAD, CCS/HIFI, ISOSEQ, and UNROLLED — which set appropriate minimap2 scoring/parameters for each data type. A two-stage workflow is common: first building an index from the reference, then aligning reads against that index with sorting enabled. Because it emits sorted PacBio-compatible BAMs with correct tags, its output feeds cleanly into downstream PacBio tools such as Sniffles, DeepVariant, and pbsv. | Long Read Aligner, Graph-Guided Mapping, Reference Genome Mapping, Bioinformatics | Bioinformatics, Biological Sciences | Alignment Tool | Documentation, Uses, and more |
| pbtk | lcc | 1.0 | pbtk (PacBio BAM toolkit), version 3.1.1, is a collection of small command-line utilities for working with the BAM files that PacBio sequencers produce. Because PacBio delivers reads (subreads, CCS/HiFi reads) in BAM rather than FASTQ, these tools bridge PacBio data into standard downstream pipelines: bam2fastq and bam2fasta convert PacBio BAM records into gzipped FASTQ or FASTA (preserving read names and grouping), while pbindex builds the .pbi index that many PacBio tools require, and pbmerge combines multiple BAM files. Inputs are PacBio BAM files (optionally with a dataset XML) and outputs are the converted sequence files or index. It is typically used to obtain FASTQ from HiFi read BAMs for assembly or mapping. pbtk is a lightweight but essential glue utility in any PacBio HiFi/CLR long-read workflow. | Documentation, Uses, and more | |||
| pcangsd | mcc | 1.0 | PCAngsd (Principal Component Analysis of next-generation sequencing data) is a Python/Cython tool for population-genetic inference from genotype likelihoods rather than called genotypes, making it well-suited to low- and medium-coverage sequencing where hard genotype calls are unreliable. Taking Beagle-format genotype-likelihood files (typically produced by ANGSD), it iteratively estimates individual allele frequencies and a covariance matrix whose eigen-decomposition yields principal components describing population structure. Beyond PCA it can estimate admixture proportions, per-site inbreeding coefficients and Hardy-Weinberg departures, kinship, selection statistics, and can call genotypes conditioned on the inferred structure, producing a covariance matrix for downstream eigen-analysis. This deployment provides PCAngsd (two builds, around v0.98) and is a standard tool in low-coverage population-genomics pipelines alongside ANGSD. | Population Genetics, Genomic Data Analysis | Genetics, Biological Sciences | Tool | Documentation, Uses, and more |
| pcre | mcc, lcc | 1.0 | The PCRE library is a set of functions that implement regular expression pattern matching using the same syntax and semantics as Perl 5. NOTE:This version of PCRE is now at end of life, and is no longer being actively maintained. New Projects should use PCRE2 instead. Description Source: https://www.pcre.org/ |
Regular Expressions, Pattern Matching, Library | Biology, Computer & Information Sciences | Runtime Library | Documentation, Uses, and more |
| pcre2 | mcc, ecc | N/A | PCRE2 is a set of functions that implement regular expression pattern matching using the same syntax and semantics as Perl 5. PCRE has its own native API, as well as a set of wrapper functions that correspond to the POSIX regular expression API. NOTE: This is the current version of the PCRE library. Description Source: https://www.pcre.org/ |
Regular Expressions, Text Processing, Pattern Matching | Biology, Computer & Information Sciences | C Library | Documentation, Uses, and more |
| pdftotext | lcc | N/A | pdftotext is a command-line utility that converts PDF documents into plain text files, allowing for easier text extraction and manipulation. | PDF, Text Extraction, Command Line, Open Source | Information Retrieval, Computer Science | Utility | Documentation, Uses, and more |
| pdtoolkit | mcc, lcc | N/A | PDT is a framework for analyzing source code | Static Analysis, Source Code Parsing | Development library | Documentation, Uses, and more | |
| pegtl | lcc | N/A | PEGTL (Parsing Expression Grammar Template Library) is a C++ library designed for parsing expression grammars, providing a framework for implementing parsers in a straightforward manner. | C++, Parsing, Grammar, Template Library | Software Engineering, Other Computer and Information Sciences | Library | Documentation, Uses, and more |
| perl | mcc, lcc, ecc | N/A | Perl is a highly capable, feature-rich programming language with over 36 years of development. Perl runs on over 100 platforms from portables to mainframes and is suitable for both rapid prototyping and large scale development projects. Description Source: https://www.perl.org/about.html |
Programming Language, Text Processing, Scripting, Web Development, System Administration | Bioinformatics, Computer & Information Sciences | Programming Language | Documentation, Uses, and more |
| perl-bio-viennangs | mcc | 1.0 | Bio::ViennaNGS is a Perl library (a distribution of modules and accompanying scripts) that provides reusable building blocks for constructing next-generation sequencing data-analysis workflows, developed by the ViennaRNA/TBI group. It offers object-oriented modules for handling common NGS data structures and file formats (BED, BAM/SAM via samtools, BigWig/BigBed) and utility scripts for tasks such as generating coverage tracks, extracting and manipulating genomic feature annotations, computing feature overlaps and statistics, and preparing data for genome-browser visualization. It is intended to be composed into custom pipelines rather than run as a single application, exposing functionality through its Perl API and bundled command-line scripts. This is version 0.19.2. | Documentation, Uses, and more | |||
| perl-bioperl | mcc, lcc | 1.0 | BioPerl (version 1.7.2 from bioconda) is a long-standing collection of Perl modules providing reusable building blocks for bioinformatics programming. Its core capabilities include reading and writing a wide range of sequence and alignment formats through Bio::SeqIO and Bio::AlignIO, manipulating sequence and feature objects, parsing the output of common tools (BLAST, HMMER, and others via Bio::SearchIO), accessing remote databases, and working with annotations and phylogenetic trees. Rather than being a single command, it is a library that scripts import to compose custom pipelines and format conversions. It remains a dependency for many established genomics tools and legacy pipelines, and is frequently pulled in specifically to satisfy those Perl-based workflows. In this container it is provided in its own conda environment exposed as a Singularity app so BioPerl-dependent scripts have a consistent runtime. | bioinformatics, perl, sequence analysis, genomics, molecular biology | Computational Biology, Biological Sciences | Library | Documentation, Uses, and more |
| perl-data-dumper | mcc, lcc | N/A | Documentation, Uses, and more | ||||
| perl-encode-locale | mcc | N/A | A Perl module that provides a way to determine the locale encoding of the environment and to set the encoding for input and output streams accordingly. | Perl, Encoding, Locale, Internationalization | Natural Language Processing, Computer Science | Library | Documentation, Uses, and more |
| perl-extutils-config | mcc | N/A | perl-extutils-config is a Perl module that provides a way to retrieve configuration information about Perl installations and modules. | Perl, Configuration, Build Tools | Software Engineering, Other Computer and Information Sciences | Library | Documentation, Uses, and more |
| perl-extutils-helpers | mcc | N/A | Documentation, Uses, and more | ||||
| perl-extutils-installpaths | mcc | N/A | A Perl module that provides methods to determine installation paths for Perl modules and scripts. | Perl, Installation, Module Management | Software Engineering, Computer Science | Library | Documentation, Uses, and more |
| perl-file-listing | mcc | N/A | A Perl module that provides a way to create directory listings in a structured format, often used for generating HTML file listings. | Perl, File Management, Web Development | Software Engineering, Other Computer and Information Sciences | Library | Documentation, Uses, and more |
| perl-html-parser | mcc | N/A | A Perl module for parsing HTML documents and extracting data from them. | HTML, Parsing, Perl, Web Scraping | Computer Science, Software Engineering | Library | Documentation, Uses, and more |
| perl-html-tagset | mcc | N/A | The HTML::Tagset module provides a set of constants for HTML tags and attributes, which can be used in conjunction with HTML::Parser and other modules to facilitate HTML parsing and manipulation. | a, div, span, p, img, table, tr, td | Computer Science, Software Engineering | Library | Documentation, Uses, and more |
| perl-http-cookies | mcc | N/A | A Perl module for managing HTTP cookies, allowing for easy handling of cookie storage and retrieval in web applications. | Perl, HTTP, Cookies, Web Development | Software Engineering, Computer Science | Library | Documentation, Uses, and more |
| perl-http-daemon | mcc | N/A | A simple HTTP server written in Perl that can be used for testing and development purposes. | Perl, HTTP, Web Server, Development | Computer Science, Software Engineering | Web Server | Documentation, Uses, and more |
| perl-http-date | mcc | N/A | perl-http-date is a Perl module that provides functions for parsing and formatting HTTP date strings. | Perl, HTTP, Date, Parsing, Formatting | Software Engineering, Computer Science | Library | Documentation, Uses, and more |
| perl-http-message | mcc | N/A | perl-http-message is a Perl module that provides classes for HTTP message handling, including HTTP requests and responses. | Perl, HTTP, Web Development, Networking | Computer Science | Library | Documentation, Uses, and more |
| perl-http-negotiate | mcc | N/A | A Perl module that provides content negotiation features for HTTP requests. | Perl, HTTP, Content Negotiation, Web Development | Computer Science, Software Engineering | Library | Documentation, Uses, and more |
| perl-io-html | mcc | N/A | The IO::HTML module provides an interface for reading HTML documents as if they were plain text files, allowing for easier parsing and manipulation of HTML content. | HTML, Perl, Parsing, Text Processing | Software Engineering, Informatics | Library | Documentation, Uses, and more |
| perl-libwww-perl | mcc | N/A | A collection of Perl modules that provide a simple and consistent interface for web programming. | Perl, Web Development, HTTP Client, Networking | Software Engineering, Other Computer and Information Sciences | Library | Documentation, Uses, and more |
| perl-lwp-mediatypes | mcc | N/A | A Perl module that provides a way to handle media types (MIME types) in web applications. | Perl, Web Development, MIME Types, HTTP | Software Engineering, Applied Computer Science | Library | Documentation, Uses, and more |
| perl-module-build | mcc | N/A | Perl Module Build is a system for building and installing Perl modules. It provides a simple way to create a Makefile for your Perl module, allowing for easy installation and management of dependencies. | Perl, Build System, Module Management | Computer Science, Software Engineering | Build Tool | Documentation, Uses, and more |
| perl-module-build-tiny | mcc | N/A | A minimalistic Perl module for building and installing Perl modules. | Perl, Build Tools, Module Management | Software Engineering | Library | Documentation, Uses, and more |
| perl-net-http | mcc | N/A | perl-net-http is a Perl library that provides an interface for HTTP client functionality, allowing users to send HTTP requests and handle responses. | Perl, HTTP, Networking, Web Development | Computer Science | Library | Documentation, Uses, and more |
| perl-test-needs | mcc | N/A | Documentation, Uses, and more | ||||
| perl-text-soundex | mcc | N/A | A Perl module that implements the Soundex phonetic algorithm for indexing names by sound, as pronounced in English. | Phonetics, Text Processing, Perl | Computational Linguistics, Other Natural Sciences | Library | Documentation, Uses, and more |
| perl-try-tiny | mcc | N/A | Try::Tiny is a minimalistic module for exception handling in Perl, providing a simple way to catch exceptions without the overhead of traditional Perl eval blocks. | Perl, Exception Handling, Error Management, Lightweight, Programming | Software Engineering, Other Computer and Information Sciences | Library | Documentation, Uses, and more |
| perl-uri | mcc | N/A | perl-uri is a Perl module for manipulating Uniform Resource Identifiers (URIs). It provides a simple way to create, parse, and manipulate URIs in a consistent manner. | Perl, URI, Web Development, Networking | Software Engineering, Computer Science | Library | Documentation, Uses, and more |
| perl-velvetoptimiser | mcc, lcc | 1.0 | VelvetOptimiser (version 2.2.6) is a Perl wrapper that automates parameter optimization for the Velvet de novo short-read assembler. Choosing a good hash length (k-mer size) and coverage cutoff strongly affects Velvet's assembly quality, and VelvetOptimiser systematically sweeps a user-defined range of k values, running the velveth/velvetg stages at each, and selects the best assembly according to an optimization metric — by default maximizing N50 for the k-mer search and using a cutoff estimation for the coverage/expected-coverage parameters, with the criteria customizable. This removes the tedious manual trial-and-error of tuning Velvet by hand. Inputs are the sequencing reads (in Velvet-supported formats) and the k-mer range/step; output is the optimized assembly directory plus a log of the parameter search. It targets bacterial and small-genome short-read assembly projects where Velvet is used. | Documentation, Uses, and more | |||
| perl-www-robotrules | mcc | N/A | Documentation, Uses, and more | ||||
| perl-xml-parser | mcc | N/A | A Perl module for parsing XML documents. | XML, Parsing, Perl, SAX, DOM | Software Engineering, Computer Science | Library | Documentation, Uses, and more |
| petsc | mcc, lcc | N/A | Portable Extensible Toolkit for Scientific Computation | Scientific Computing, Computational Science, Numerical Analysis | Applied Mathematics, Natural Sciences | Library | Documentation, Uses, and more |
| pggb | mcc | 1.0 | PGGB (the PanGenome Graph Builder, version 0.5.4 from bioconda) constructs unbiased, reference-free pangenome variation graphs from a collection of input genome assemblies, capturing SNPs, indels, and structural variation shared across the set. The pipeline first performs all-versus-all alignment of the input sequences with wfmash, induces a graph with seqwish, then progressively normalizes and simplifies the graph topology with smoothxg and GFAffix. Its inputs are a single bgzipped, samtools-faidx-indexed FASTA containing all genomes, and key parameters include the number of haplotypes, a mapping identity threshold, and a segment length; outputs are a GFA graph plus ODGI visualizations and optionally a VCF of variants relative to a chosen path. It is a cornerstone of the PanGenome Graphs (PGGB/ODGI/vg) toolchain used in the Human Pangenome Reference Consortium. | Performance Evaluation, Benchmarking, Computational Applications | Performance Evaluation & Benchmarking, Training | Documentation, Uses, and more | |
| pgi | mcc, lcc | N/A | The NVIDIA HPC Software Development Kit (SDK) includes the proven compilers, libraries and software tools essential to maximizing developer productivity and the performance and portability of HPC applications. Description Source: https://developer.nvidia.com/hpc-sdk |
Hpc, Compilers, Parallel Programming | Computer Science, Engineering & Technology | Development Tool | Documentation, Uses, and more |
| pgi-cuda | lcc | N/A | Documentation, Uses, and more | ||||
| pgs-recon | ecc | 1.0 | PGS-Recon is an Open OnDemand (OOD) interactive application registered on the ECC cluster for browser-based photogrammetry reconstruction — building 3D models and point clouds from sets of overlapping 2D photographs. Photogrammetry pipelines detect and match features across many images, estimate camera positions via structure-from-motion, and then generate a dense point cloud, mesh, and textured 3D model that can be used in research areas such as cultural-heritage documentation, geosciences, biology, and engineering. As an OOD interactive app it launches from the cluster's web portal, provisioning compute resources so users upload image sets and run the reconstruction through a GUI rather than the command line. In the SDS catalog this entry is a catalog-only stub that registers the app's existence and access path rather than a built container image; it points researchers to the ECC Open OnDemand interface for running photogrammetry jobs. | Documentation, Uses, and more | |||
| phasebook | mcc | 1.0 | phasebook is a de novo, reference-free assembler that reconstructs both haplotypes of a diploid genome directly from long reads, producing haplotype-resolved (phased) contigs without needing a reference genome for phasing. It clusters reads by haplotype based on their overlap and heterozygous-variant patterns and then assembles each haplotype separately, so structural and single-nucleotide differences between the two parental copies are preserved rather than collapsed. It works with PacBio and Oxford Nanopore data and, in this environment, relies on a set of bundled tools — minimap2 for overlaps, WhatsHap and Longshot for variant/phasing information, samtools/bcftools, and racon and fpa — to carry out its pipeline. Input is a long-read FASTQ/FASTA dataset plus configuration of the sequencing platform; output is a set of phased haplotype-resolved contigs. It is installed from the phasebook GitHub repository. | Documentation, Uses, and more | |||
| phaser | lcc | N/A | PHASER is a software package for macromolecular crystallography that is used for solving crystal structures using molecular replacement and experimental phasing techniques. | Crystallography, Structural Biology, Computational Chemistry | Biochemistry, Molecular Biology | Application | Documentation, Uses, and more |
| phastest | mcc | 1.0 | PHASTEST (PHAge Search Tool with Enhanced Sequence Translation) is the latest generation of the PHAST/PHASTER family of tools for rapidly identifying, annotating, and scoring prophage (integrated bacteriophage) regions within bacterial genome sequences. Given a bacterial genome as a FASTA nucleotide sequence or a GenBank annotation, it predicts genes, compares predicted proteins against phage- and virus-specific databases, and identifies clusters of phage-like genes, then classifies each detected prophage region as intact, questionable, or incomplete based on the density and identity of recognizable phage genes (integrases, capsid/tail proteins, terminases, etc.) and attachment sites. It bundles the supporting bioinformatics dependencies it needs — BLAST+, DIAMOND, Prodigal for gene calling, tRNAscan-SE and Aragorn for tRNA detection, and Barrnap for rRNA — so it runs as a self-contained prophage-annotation pipeline. It is commonly used in microbial genomics to catalog the mobile phage content of bacterial isolates. This deployment builds PHASTEST from the upstream application download on an Ubuntu 20.04 base. | Documentation, Uses, and more | |||
| phdf5 | lcc | N/A | A general purpose library and file format for storing scientific data | Parallel I/O, Phdf5, Hdf5, Large Datasets | Computer & Information Sciences | File Format | Documentation, Uses, and more |
| phrynomics | mcc | 1.0 | Phrynomics is an R package for handling and analyzing SNP datasets in phylogenetic and population-genetic studies, developed to streamline the preparation and filtering of reduced-representation (e.g. RADseq-style) SNP data for downstream tree inference. It provides functions to read SNP matrices, translate between coding schemes, filter sites (for example removing invariant or non-binary sites and managing missing data), and export the data into the input formats required by phylogenetics programs, helping researchers move from raw SNP tables to analysis-ready alignments. It is installed from GitHub (bbanbury/phrynomics via remotes) and provided within a Rocky 9.3 + Miniconda R 4.4 container built for population genetics that also bundles dartRverse and TESS3r along with their geospatial and compiled dependencies. Typical use loads the package in R and applies its SNP-import and filtering functions before writing out data for phylogenetic estimation. | Documentation, Uses, and more | |||
| phylophlan | lcc | 1.0 | PhyloPhlAn 3 (version 3.0.2 from bioconda) is a tool for large-scale microbial phylogenetic profiling and phylogeny reconstruction from genomes and metagenome-assembled genomes (MAGs). It can place unknown genomes/MAGs into a species-level context, assign them to known species genome bins (SGBs), and build high-resolution phylogenies ranging from strain-level trees within a species to trees spanning the whole microbial tree of life, by extracting and aligning conserved marker genes and then inferring a tree. It orchestrates external aligners and tree builders (e.g., Diamond/USEARCH for marker search, MAFFT/MUSCLE, trimAl, and RAxML/IQ-TREE/FastTree) through configurable pipeline presets, and provides utilities to generate the marker/config databases. Given a directory of genomes, a chosen database, and a diversity setting, it produces a multiple-sequence alignment and a Newick tree. It is widely used in comparative and metagenomic microbiology for strain tracking and taxonomic placement. | bioinformatics, phylogenetics, genomics, microbiology | Microbial phylogenetics, Biological Sciences | Command-line tool | Documentation, Uses, and more |
| phyluce | lcc | N/A | Phyluce is a software package designed for the analysis of phylogenomic data, particularly for handling and processing genomic data for phylogenetic studies. | phylogenetics, bioinformatics, genomics, sequence analysis | Genomics, Phylogenetics | Analysis | Documentation, Uses, and more |
| phyml | lcc | 1.0 | PhyML (version 3.3.20200621) estimates maximum-likelihood phylogenetic trees from nucleotide or amino-acid multiple sequence alignments. It implements fast hill-climbing topology search (including SPR and NNI moves) and supports a wide range of substitution models (e.g. HKY85, GTR for DNA; LG, WAG, JTT for proteins) with rate heterogeneity across sites (gamma categories and invariant sites), plus model selection via SMS. Branch-support can be assessed with bootstrap replicates or faster approximate likelihood-ratio tests (aLRT/SH-like and aBayes). Input is a PHYLIP-format alignment; outputs are the inferred tree(s) in Newick along with a statistics file reporting the likelihood, model parameters, and support values. It can be run interactively via a menu or non-interactively with command-line flags. In this environment the conda env also includes UCX for high-performance interconnect support. PhyML is a long-standing, well-cited option for ML phylogenetics. | phylogenetics,bioinformatics,maximum likelihood,tree estimation | Bioinformatics, Phylogenetics | Standalone | Documentation, Uses, and more |
| picard | mcc, lcc | 1.0 | Picard (this build uses picard.jar version 2.26.4, run on Java 8) is a set of Java command-line tools from the Broad Institute for manipulating high-throughput sequencing data and its associated file formats (SAM/BAM/CRAM, VCF/BCF, and interval lists). It bundles dozens of utilities addressing common preprocessing and QC needs: MarkDuplicates (flagging or removing PCR/optical duplicate reads), SortSam, AddOrReplaceReadGroups, MergeSamFiles, CreateSequenceDictionary, BedToIntervalList, and a broad family of metrics collectors (CollectAlignmentSummaryMetrics, CollectInsertSizeMetrics, CollectHsMetrics, etc.). Picard is a standard component of the GATK Best Practices pipeline, most notably for duplicate marking and for building the reference sequence dictionary and read groups required by downstream variant callers. It is a ubiquitous utility in resequencing and variant-analysis workflows. | Bioinformatics, Ngs, Hts, Sequencing Data, Sam, Bam, Vcf | Genomics, Biological Sciences | Command Line Tool | Documentation, Uses, and more |
| picrust2 | mcc | 1.0 | PICRUSt2 (version 2.5.2 from bioconda) predicts the functional gene content of microbial communities from marker-gene amplicon data (typically 16S rRNA, but also 18S/ITS), inferring metagenome function without shotgun sequencing. It places study sequences (ASVs) onto a reference phylogeny with EPA-ng/gappa, uses hidden-state prediction (castor) to infer per-organism gene family copy numbers, and then combines these with per-sample abundances to output predicted KO, EC, and MetaCyc pathway abundances, along with NSTI values that flag poorly characterized sequences. The standard workflow is a single wrapper pipeline that takes a sequence FASTA and a feature (BIOM) table and runs placement, hidden-state prediction, metagenome prediction, and pathway inference end to end. Its outputs integrate readily with downstream differential-abundance and pathway-enrichment analysis. It is widely used in microbiome studies to generate functional hypotheses from inexpensive amplicon data. | Bioinformatics, Metagenomics, Microbiome Analysis, Functional Gene Prediction | Biological Sciences | Metagenomics Tool | Documentation, Uses, and more |
| pigz | lcc, ecc | N/A | pigz, which stands for parallel implementation of gzip, is a fully functional replacement for gzip that exploits multiple processors and multiple cores to the hilt when compressing data. pigz was written by Mark Adler, and uses the zlib and pthread libraries. Description Source: https://zlib.net/pigz/ |
Compression, Parallel Processing, Utility, Command-Line Tool | Computer Science, Computer & Information Sciences | Utility | Documentation, Uses, and more |
| pilon | lcc | 1.0 | Pilon (bioconda 1.23) is an automated tool for improving draft genome assemblies and for variant detection in small genomes, using read alignments to identify and correct assembly errors. Given a reference/draft assembly FASTA and one or more BAM files of reads mapped to it, Pilon analyzes the pileup and mate-pair information to correct single-base errors and small indels, fill gaps, and resolve some local misassemblies, and it can also report variants (SNPs and indels) relative to the input. It works well combining Illumina short-read evidence to polish long-read assemblies, and it can process multiple libraries (fragment, jump, and unpaired) to inform its corrections. Inputs are the assembly FASTA plus sorted, indexed BAM alignments; outputs are a corrected FASTA (with '_pilon' suffixed contig names), an optional VCF of changes, and change/track files documenting every modification. It is a standard polishing step, often run iteratively, in microbial and small-eukaryote assembly pipelines. | Genome Assembly, Genomics, Bioinformatics, Computational Biology | Genomics, Biological Sciences | Tool | Documentation, Uses, and more |
| pipeline | mcc | 1.0 | This entry is a single integrated conda environment ('pipeline-1') that bundles a complete, ready-to-use RNA-seq and variant-calling toolchain so that a full analysis can be run without switching environments. For alignment it provides classic and modern read mappers: TopHat and Bowtie2, the splice-aware RNA-seq aligners STAR and HISAT2, and the DNA aligner BWA-MEM2. For transcript assembly and quantification it includes Cufflinks and StringTie. For file manipulation and variant workflows it bundles SAMtools, BCFtools, VCFtools, and the variant/QC pair GATK4 and Picard, plus the BioPerl and Biopython libraries for scripting. In practice a user can align reads (STAR/HISAT2 for RNA, BWA-MEM2 for DNA), assemble/quantify transcripts (StringTie), and call and filter variants (GATK4/bcftools) entirely within this one environment. It targets standard transcriptomics and resequencing projects where a curated, version-consistent set of common tools is more convenient than assembling them individually. | Documentation, Uses, and more | |||
| pixman | mcc | N/A | Pixman is a low-level software library for pixel manipulation, providing features such as image compositing and trapezoid rasterization. Important users of pixman are the cairo graphics library and the X server. Description Source: https://www.pixman.org/ |
Pixel Manipulation, Rendering, Software Library, Performance | Software Development, Computer & Information Sciences | Graphics Library | Documentation, Uses, and more |
| pkgconf | mcc, lcc, ecc | N/A | pkgconf is a program which helps to configure compiler and linker flags for development frameworks. It is similar to pkg-config from freedesktop.org, providing additional functionality while also maintaining compatibility. Description Source: http://pkgconf.org/ |
Package Configuration, Build Flags, Link Flags, System-Agnostic | Software Engineering, Computer & Information Sciences | Package Configuration | Documentation, Uses, and more |
| plascad | lcc | 1.0 | Plascad is a computationally efficient tool for classifying plasmid sequences by their mobility and annotating the antibiotic-resistance genes (ARGs) they carry—both central to understanding the horizontal spread of resistance and virulence between bacterial cells. Working from plasmid nucleotide sequences, it first predicts open reading frames, then uses profile-HMM searches (hmmsearch against built HMM protein profiles) to detect the protein machinery of DNA transfer—relaxase (MOB), the type IV coupling protein (T4CP), and type IV secretion system (T4SS/mate-pair-formation) components—and on that basis classifies each plasmid as conjugative (relaxase + T4CP + T4SS), mobilizable (relaxase only), or non-mobilizable. It additionally identifies ARGs by searching against a structured resistance-gene database and can render annotated plasmid maps for visualization. Typical usage takes a multi-FASTA of plasmid sequences and outputs per-plasmid mobility classifications and annotation tables/figures. It is used in microbial genomics and antimicrobial-resistance surveillance to assess the horizontal-transfer potential of plasmids recovered from isolates or metagenomes. | Documentation, Uses, and more | |||
| plasma | lcc | N/A | Parallel Linear Algebra Software for Multicore Architectures | Documentation, Uses, and more | |||
| plink | mcc, lcc | 1.0 | PLINK is a comprehensive, widely used toolset for whole-genome and genome-wide association study (GWAS) analysis of genotype/phenotype data, developed originally for human population genetics but applied across many species. It handles data management and format conversion (PED/MAP, binary BED/BIM/FAM, VCF), extensive quality control (missingness, minor-allele-frequency, Hardy-Weinberg equilibrium filtering), and population-genetic computations including allele/genotype frequencies, linkage-disequilibrium estimation and LD-based pruning, identity-by-descent/relatedness, PCA, and Mendel-error checks. Its association testing covers basic and covariate-adjusted case/control and quantitative-trait models. This is the stable version 1.90b6.21 (PLINK 1.9). | Genetic Analysis, Genomic Data Analysis, Gwas, Genetic Epidemiology | Genetics, Biological Sciences | Tool | Documentation, Uses, and more |
| plink2 | mcc | 1.0 | PLINK 2.0 is a comprehensive, high-performance toolset for whole-genome and large-scale genotype/phenotype data management and statistical genetics analysis, the successor to the classic PLINK 1.9 with a redesigned file format and greatly improved speed and memory efficiency for biobank-scale datasets. It reads and writes its native binary formats (the pgen/pvar/psam trio, as well as legacy bed/bim/fam and VCF/BCF) and performs quality control filtering (by missingness, minor-allele frequency, Hardy-Weinberg equilibrium), sample and variant management, LD-based pruning, principal-component analysis, kinship/relatedness estimation (KING-robust), and genome-wide association testing via linear and logistic regression with covariates (its --glm association mode), including dosage/imputed-data support. It is a core workhorse in human and non-human GWAS and population-genetics pipelines. This build is version 2.00a5.12 from Bioconda. | Genomics, Association Analysis, Genotype-Phenotype Data | Genetics | Tool | Documentation, Uses, and more |
| plotsr | mcc | 1.0 | plotsr is a visualization tool for comparative genomics that draws clear, publication-quality diagrams of structural rearrangements and local sequence variations between multiple chromosome-level genome assemblies arranged in a stack. It takes as input the annotated structural-variation output from SyRI (Synteny and Rearrangement Identifier) — which classifies syntenic regions, inversions, translocations, and duplications between assembly pairs — together with the genome FASTA files (and their .fai indices), and renders the syntenic and rearranged blocks as colored ribbons connecting the chromosomes of successive genomes, so that inversions, translocations, and duplications are immediately visible along the genome. Users can chain several genomes to visualize rearrangements across a series of assemblies, add tracks (e.g. gene density, GC content) via a bedgraph/track file, and mark specific regions of interest. It is commonly the final visualization step after a minimap2-then-SyRI genome-comparison pipeline. This build is version 1.1.1 from Bioconda. | Plotting, Data Visualization, Python Library | Computer & Information Sciences | Data Visualization | Documentation, Uses, and more |
| plumed | lcc | N/A | PLUMED is an open-source library implementing enhanced-sampling algorithms, various free-energy methods, and analysis tools for molecular dynamics simulations. Description Source: https://www.plumed.org/ |
Molecular Dynamics, Enhanced Sampling, Free Energy Calculations, Collective Variables | Chemistry, Biological Sciences | Library | Documentation, Uses, and more |
| pmix | mcc, lcc, ecc | N/A | The Process Management Interface (PMI) has been used for quite some time as a means of exchanging wireup information needed for inter-process communication. Two versions (PMI-1 and PMI-2) have been released as part of the MPICH effort, with PMI-2 demonstrating better scaling properties than its PMI-1 predecessor. Description Source: https://pmix.github.io/standard |
Process Management, Parallel Computing, High-Performance Computing, Data Exchange, Resource Management, Scalability | Computer Science, Engineering & Technology | Hpc Tool | Documentation, Uses, and more |
| pnetcdf | lcc | N/A | A Parallel NetCDF library (PnetCDF) | Parallel I/O, Netcdf Files, Scalability, Performance | High-Performance Computing, Physical Sciences | I/O Library | Documentation, Uses, and more |
| popoolation | mcc | 1.0 | PoPoolation is a collection of Perl and R scripts for population-genetics analysis of pooled next-generation sequencing (Pool-seq) data from a single population, where many individuals are sequenced together as one pooled sample. It computes classic diversity and neutrality statistics — Tajima's Pi (nucleotide diversity), Watterson's Theta, and Tajima's D — either genome-wide in sliding windows or for defined genic regions, correcting for the biases introduced by pooling and finite pool size. The workflow starts from a mpileup file produced by samtools from reads aligned with bwa, applies subsampling and quality/coverage filtering, and outputs per-window statistics plus files suitable for genome-browser visualization; it also includes utilities for indel filtering and gene-based (nonsynonymous vs synonymous) analyses. This is PoPoolation 1.2.2 (with PoPoolation2 for pairwise/population-differentiation Fst analyses also available). | bioinformatics, population genetics, next-generation sequencing | Population Genetics, Genomics | Analysis Tool | Documentation, Uses, and more |
| popt | mcc | N/A | Popt is a command line option parsing library that provides a simple, yet powerful mechanism for parsing command-line options and arguments. It assists in simplifying the process of writing command-line interfaces for applications. | Command Line Options, C Library | Computer & Information Sciences | Command Line Option Parser | Documentation, Uses, and more |
| porechop | mcc | 1.0 | Porechop is an adapter-finding and trimming tool for Oxford Nanopore sequencing reads. It searches each read for the known set of ONT adapter sequences at both ends and removes them, and — importantly — it also detects adapters located in the middle of reads, which indicate chimeric molecules where two DNA fragments were sequenced as one; such reads are split at the internal adapter. Porechop can additionally demultiplex barcoded runs by recognizing barcode-specific adapters and binning reads accordingly. It takes FASTQ (or FASTA, gzip-supported) input and writes trimmed reads, with options to discard rather than split chimeras and controls over adapter-match thresholds. Note the tool is no longer actively maintained upstream but remains widely used; this is version 0.2.4 from bioconda. | bioinformatics, sequence analysis, data processing | Bioinformatics, Genomics | Command-line tool | Documentation, Uses, and more |
| prank | lcc | 1.0 | PRANK (version v.150803 from bioconda) is a probabilistic, phylogeny-aware progressive multiple sequence aligner for DNA, codon, and amino-acid data. Its distinguishing feature is that it uses an evolutionary (guide) tree to place insertions and deletions correctly and, crucially, avoids over-aligning independent insertions by treating insertions and deletions differently, which reduces the systematic errors that inflate downstream inferences of positive selection. It can align coding sequences in codon-aware mode, output ancestral sequences, and produce alignments annotated with inferred indel events (in extended formats such as HSAML/XML), yielding a FASTA alignment and, in structure-aware modes, per-site event information. It is favored in molecular-evolution studies where alignment quality strongly affects dN/dS and phylogenetic conclusions. | Sequence Alignment, Bioinformatics, Genomics, Phylogenetics | Genetics, Biological Sciences | Computational Software | Documentation, Uses, and more |
| pretext-suite | mcc, lcc | 1.0 | PretextSuite is a set of high-performance tools for generating and interactively visualizing Hi-C contact maps, used primarily in the manual curation and scaffolding stages of chromosome-level genome assembly. It comprises PretextMap, which ingests alignment data (SAM/BAM of Hi-C reads mapped to an assembly, typically piped from samtools) and builds a compact, texture-based .pretext contact-map file; PretextView, a fast GPU-accelerated GUI for panning, zooming, and editing the contact map to detect misassemblies, order/orient scaffolds, and assign chromosomes; and PretextGraph, which overlays additional 1D signal tracks (such as coverage or gaps) onto the map. In a typical pipeline Hi-C alignments are streamed into PretextMap to build the contact-map file, which is then opened in PretextView for curation. The Hi-C signal reveals the physical proximity of sequences, making structural errors visible as off-diagonal patterns. This is bioconda pretext-suite 0.0.2. | Documentation, Uses, and more | |||
| prodigal-2.6.3 | mcc | 1.0 | Prodigal (Prokaryotic Dynamic Programming Genefinding Algorithm), version 2.6.3, is a fast, unsupervised gene-prediction tool for identifying protein-coding genes in bacterial, archaeal, and viral genomes. It learns the coding characteristics of a genome directly from the input sequence (GC bias, ribosomal binding-site motifs, start-codon usage) with no need for training data, then predicts complete and partial genes along with accurate translation initiation (start) sites. It handles finished genomes in single mode and fragmented assemblies/metagenomic contigs in an anonymous/meta mode where per-contig training is impractical. It emits predicted protein and nucleotide sequences plus coordinate output in GFF, GenBank, or Sequin table formats. It is a foundational step in prokaryotic annotation pipelines (e.g. Prokka, DRAM) and in binning/metagenomics workflows. | Documentation, Uses, and more | |||
| prokka | lcc | 1.0 | Prokka is a pipeline for rapid, standardized annotation of prokaryotic genomes — bacteria, archaea, and viruses — producing publication- and submission-ready output in minutes on a typical microbial assembly. Given an assembled contigs FASTA, Prokka first predicts coding sequences with Prodigal, then identifies tRNAs (Aragorn), rRNAs (Barrnap), and other features, and assigns functional annotations by hierarchically searching curated databases (its bundled BLAST+/HMMER against UniProt, Pfam/TIGRFAM HMMs, and optional genus-specific databases). It outputs a full set of coordinated files — GFF3, GenBank (.gbk), EMBL, annotated FASTA of genes/proteins (.ffn/.faa), and a feature table — with consistent locus tags suitable for direct submission to GenBank/ENA. It can optionally be pointed at a trusted reference protein set to prioritize annotation and at a specific genus to tune naming. It remains a standard first annotation step for bacterial genome projects. This is version 1.14.6 from Bioconda. | Genome Annotation, Bioinformatics, Genomics | Genetics, Biological Sciences | Annotation Tool | Documentation, Uses, and more |
| proovread | lcc | N/A | proovread is a hybrid error-correction pipeline for long, error-prone third-generation sequencing reads (primarily PacBio, with Nanopore support), provided on LCC as a conda module (ccs/conda/proovread-2.14.0). It uses high-accuracy Illumina short reads and iterative short-read consensus to correct long reads, substantially lowering their error rate to improve downstream genome assembly and structural analysis. It is a real user-facing bioinformatics application (BioInf-Wuerzburg/proovread). | bioinformatics, genomics, sequence analysis, error correction | Computational Biology, Bioinformatics | Open Source | Documentation, Uses, and more |
| protobuf | lcc | N/A | Protocol Buffers are language-neutral, platform-neutral extensible mechanisms for serializing structured data. Description Source: https://protobuf.dev/ |
Serialization, Data Interchange, Code Generation | Software Engineering, Computer & Information Sciences | Library | Documentation, Uses, and more |
| prun | mcc, lcc, ecc | N/A | job launch utility for multiple MPI families | Parallel Computing, High-Performance Computing, Distributed Computing, Data Analysis, Computation | Other Computer & Information Sciences | Computational Software | Documentation, Uses, and more |
| psi4 | mcc, lcc | 1.0 | Psi4 (version 1.6.1) is an open-source quantum-chemistry package for high-accuracy ab initio electronic-structure calculations of molecular energies and properties. It implements a broad hierarchy of methods — Hartree-Fock and density-functional theory (DFT) with many functionals, second-order perturbation theory (MP2), coupled-cluster (including CCSD and CCSD(T)), symmetry-adapted perturbation theory (SAPT) for intermolecular interaction energies, and more — and supports geometry optimization, vibrational frequencies, and various molecular properties. It is designed for both efficiency and ease of use, driven either through a simple text/Python input file or as an importable Python module that integrates with the scientific Python ecosystem. A typical input defines a molecule geometry and basis set and requests a single-point energy or geometry optimization at a chosen level of theory. Psi4 is used across computational chemistry for reaction energetics, non-covalent interactions, spectroscopy, and method development, and scales with threaded/parallel execution on HPC nodes. | Quantum Chemistry, Electronic Structure, Computational Chemistry, Open Source | Theoretical Chemistry, Quantum Chemistry | Scientific Software | Documentation, Uses, and more |
| ptscotch | lcc | N/A | Graph, mesh and hypergraph partitioning library using MPI | Graph Partitioning, Parallel Computing, High Performance Computing, Scientific Computing | Computer Science, Applied Computer Science | Library | Documentation, Uses, and more |
| pugixml | lcc | N/A | pugixml is a light-weight C++ XML processing library | Xml Processing, C++ Library, Xpath, Parsing | Software Engineering, Computer & Information Sciences | Xml Processing | Documentation, Uses, and more |
| purge_dups | mcc | 1.0 | purge_dups is a genome-assembly curation tool (bioconda build 1.2.6) that identifies and removes haplotypic duplications and artefactual overlapping contigs from draft assemblies, a common problem in long-read assemblies of heterozygous or diploid genomes. It works by computing a per-base read-depth histogram from long reads mapped back to the assembly to establish coverage cutoffs, then combines that depth information with self-self alignments of the contigs to classify sequences as haplotigs, overlaps, repeats, or junk. The pipeline is a sequence of steps: map reads with minimap2 and build coverage stats with pbcstat, determine cutoffs with calcuts, split the assembly and self-align it, flag duplicates with purge_dups, and finally extract sequences to emit the purged primary assembly and a set of removed sequences. Inputs are the assembly FASTA plus long-read alignments (PAF); outputs are a deduplicated primary FASTA and the extracted duplicate/haplotig sequences. It is routinely used between contig assembly (e.g., hifiasm) and scaffolding to improve assembly accuracy and BUSCO duplication metrics. | Bioinformatics, Genomics, DNA Sequences | Genomics, Biological Sciences | Bioinformatics Tool | Documentation, Uses, and more |
| py-cython | mcc, ecc | N/A | Py-Cython is a compiler for writing C extensions for the Python language. It allows for combining Python and C code seamlessly to improve performance and efficiency. | Compiler, Python Library, Performance Optimization | Computer & Information Sciences | Development Tool | Documentation, Uses, and more |
| py-docutils | mcc | N/A | A Python package for processing reStructuredText documents, primarily used for generating documentation. | Documentation, Python, reStructuredText | Software Engineering | Library | Documentation, Uses, and more |
| py-h5py | mcc | N/A | py-h5py is a Python interface for the HDF5 binary data format, providing efficient input/output operations for working with large datasets. | Python Library, Data Format, Input/Output Operations, Large Datasets | Data Science, Computer & Information Sciences | Interface | Documentation, Uses, and more |
| py-mako | mcc, lcc | N/A | Mako is a fast and lightweight templating engine for Python, designed to be easy to use and integrate with existing applications. | templating, Python, web development | Software Engineering, Other Computer and Information Sciences | Library | Documentation, Uses, and more |
| py-markupsafe | mcc, lcc | N/A | A library for safe string handling in Python, particularly for HTML and XML. | Python, String Handling, Web Development | Software Engineering, Other Computer and Information Sciences | Library | Documentation, Uses, and more |
| py-mpi4py | mcc | N/A | py-mpi4py is a Python interface to the Message Passing Interface (MPI) standard for parallel computing, allowing users to utilize MPI functionality within Python scripts and applications. | Python, Mpi, Parallel Computing | Computer Science, Computer & Information Sciences | Library | Documentation, Uses, and more |
| py-numpy | mcc | N/A | NumPy is a fundamental package for scientific computing in Python. It provides support for arrays, matrices, and a variety of mathematical functions to operate on these data structures. | Python, Scientific Computing, Data Analysis, Numerical Computation | Applied Mathematics, Other Mathematics | Python Library | Documentation, Uses, and more |
| py-pip | lcc, ecc | N/A | py-pip is the official installer for Python packages from the Python Package Index (PyPI). It is used to install and manage software packages written in Python. | Package Management, Python, Software Installation | Software Engineering, Computer & Information Sciences, Systems & Development | Package Management | Documentation, Uses, and more |
| py-pkgconfig | mcc | N/A | A Python package that provides a way to use pkg-config from Python code, allowing for easy retrieval of compiler and linker flags for libraries. | Python, pkg-config, build tools | Software Engineering, Computer Science | Library | Documentation, Uses, and more |
| py-pygments | mcc | N/A | Pygments is a syntax highlighting library written in Python that supports a wide range of programming languages and markup formats. | syntax highlighting, programming languages, markup formats, Python | Software Engineering, Other Computer and Information Sciences | Library | Documentation, Uses, and more |
| py-setuptools | mcc, lcc, ecc | N/A | Setuptools is a package development process library designed to facilitate development and distribution of Python projects. It provides a common approach for managing project dependencies, packaging, and installation processes. | Python, Package Development, Distribution, Dependency Management | Computer Science, Computer & Information Sciences | Package Management | Documentation, Uses, and more |
| py-wheel | mcc, lcc, ecc | N/A | This is a command line tool for manipulating Python wheel files, as defined in PEP 427. It contains the following functionality:\r \r Convert .egg archives into .whl\r \r Unpack wheel archives\r \r Repack wheel archives\r \r Add or remove tags in existing wheel archives |
Python, Package Management, Software Distribution, Build Tools, Packaging | "", Software Engineering, and Development | Library | Documentation, Uses, and more |
| py2-mpi4py | lcc | N/A | Python bindings for the Message Passing Interface (MPI) standard. | Documentation, Uses, and more | |||
| py2-numpy | lcc | N/A | NumPy array processing for numbers, strings, records and objects | Documentation, Uses, and more | |||
| py2-scipy | lcc | N/A | Scientific Tools for Python | Documentation, Uses, and more | |||
| py2.7-samtools | lcc | N/A | SAMtools (version 1.10, packaged in a Python 2.7 conda environment) - a core bioinformatics application suite for manipulating high-throughput sequencing alignments in SAM/BAM/CRAM formats. It provides sorting, indexing, mpileup/variant support, format conversion, and alignment statistics (flagstat/idxstats), and is a workhorse tool in genomics pipelines. The py2.7 prefix denotes the Python 2.7 packaging variant of this conda environment. | Documentation, Uses, and more | |||
| py3-mpi4py | lcc | N/A | Python bindings for the Message Passing Interface (MPI) standard. | MPI, Python, Parallel Computing, High Performance Computing | High-Performance Computing, Computer Science | Library | Documentation, Uses, and more |
| py3-numpy | lcc | N/A | NumPy array processing for numbers, strings, records and objects | Python, Scientific Computing, Numerical Analysis, Data Science | Mathematics, Applied Mathematics | Library | Documentation, Uses, and more |
| py3-scipy | lcc | N/A | Scientific Tools for Python | Python, Scientific Computing, Numerical Analysis, Data Science | Applied Mathematics, Computer Science | Library | Documentation, Uses, and more |
| py4vasp | mcc | 1.0 | py4vasp (conda-forge 0.11.2) is the official Python API developed by the VASP group for reading, analyzing, and plotting the output of VASP electronic-structure calculations. It reads VASP's HDF5 output (vaspout.h5) and exposes structured, high-level accessors for quantities such as total and projected density of states, band structures, forces, stresses, dielectric functions, charge/magnetization densities, and structural information, avoiding manual parsing of OUTCAR/vasprun files. It integrates with the scientific Python stack, returning data as arrays and pandas-friendly structures and providing convenient interactive plots (including 3D structure and volumetric-data visualization) suitable for use in Jupyter notebooks. Typical usage instantiates a Calculation object pointing at a run directory and then retrieves or plots results such as the density of states or band structure through its accessors. It is aimed at DFT/materials-science researchers who want reproducible, scriptable post-processing of VASP runs. | Python, VASP, Computational Chemistry, Data Analysis, Visualization | Computational Chemistry, Physical Chemistry | Library | Documentation, Uses, and more |
| pyani-0.2.13.1 | mcc | 1.0 | pyani (version 0.2.13.1) computes whole-genome Average Nucleotide Identity (ANI) between microbial genomes, a key measure for prokaryotic taxonomy and species delineation (the ~95% ANI species boundary). It implements several methods selectable at runtime — ANIm (whole-genome alignment with MUMmer/NUCmer), ANIb and ANIblastall (fragment-based BLAST comparisons), and TETRA (tetranucleotide frequency correlation) — and performs all pairwise comparisons within a set of input genomes, producing matrices of percentage identity, alignment coverage, and similarity errors together with heatmap and related graphical summaries. In this 0.2.x line it is driven by the average_nucleotide_identity.py script; inputs are a directory of genome FASTA files and outputs are TSV matrices and plots. It is a standard tool for classifying isolate genomes and MAGs and for confirming or refining species assignments. | bioinformatics, genomics, sequence analysis | Microbiology, Bioinformatics | Library | Documentation, Uses, and more |
| pyannote.audio | lcc | 1.0 | pyannote.audio is a speaker diarization and voice activity detection toolkit that answers who spoke when in a recording, and it is normally the first stage before phonetic transcription because it supplies the utterance boundaries. This build runs the speaker-diarization-community-1 pipeline with its segmentation, embedding and scoring models bundled locally, so no Hugging Face token and no network access are required. For each input it writes a standard RTTM file of speaker turns and, on request, a list of speech regions; the speaker count can be fixed or bounded, and an exclusive mode produces non-overlapping turns that are ready for cutting audio. Any format FFmpeg can decode is accepted and resampling and downmixing are automatic. | Documentation, Uses, and more | |||
| pybind11 | lcc | 1.0 | pybind11 is a lightweight header-only library that exposes C++ types in Python and vice versa, mainly to create Python bindings of existing C++ code. Its goals and syntax are similar to the excellent Boost.Python library by David Abrahams: to minimize boilerplate code in traditional extension modules by inferring type information using compile-time introspection. Description Source: https://github.com/pybind/pybind11 |
Software Development, Python Library, C++ Integration, Header-Only Library | Software Engineering, Computer & Information Sciences | Interoperability | Documentation, Uses, and more |
| pygenometracks | mcc | 1.0 | pyGenomeTracks (from the deepTools group) produces publication-quality, precisely aligned genome-browser-style figures by stacking multiple data tracks over a shared genomic coordinate range. It supports a wide variety of track types, including bigWig and bedGraph signal tracks, BED/GTF gene and feature annotations, links/arcs, Hi-C contact matrices (from hicmatrix), TADs, and vertical highlights, with per-track styling controlled through a plain-text .ini configuration file that specifies each track's file, height, color, scale, and label. The typical workflow first generates a template configuration file listing the desired track files, then edits it to taste, and finally renders a chosen genomic region to PNG, SVG, or PDF. It pulls in matplotlib, numpy, pysam, pyBigWig, hicmatrix, and gffutils to read the various formats. Version 3.9 is provided as a conda tool in a Rocky 9 long-read/genomics container. | Genomics, Data Visualization, Python, Genomic Data Analysis | Bioinformatics, Biological Sciences | Python Library | Documentation, Uses, and more |
| pymatgen | mcc | 1.0 | pymatgen (Python Materials Genomics) is a robust, open-source Python library that forms the analysis backbone of the Materials Project, supporting a wide range of computational materials-science and electronic-structure workflows. It provides flexible classes for crystal structures and molecules, symmetry analysis, generation and parsing of input/output for DFT codes (notably VASP, along with others), electronic and phonon band-structure and density-of-states analysis, phase diagrams and Pourbaix diagrams, diffusion analysis, and many structure-transformation and enumeration tools. In this environment it is augmented with enumlib (structure/derivative enumeration), bader (Bader charge analysis), RDKit (cheminformatics), the ocelot_api (organic-electronics toolkit), and Shapely (geometric operations), broadening it toward organic-semiconductor and charge-analysis workflows. It is used programmatically, for example reading a POSCAR into a Structure object, analyzing symmetry, and building a phase diagram from computed energies. This is conda-forge pymatgen version 2022.0.12. | materials science, computational chemistry, data analysis, python | Computational Materials Science, Materials Science | Library | Documentation, Uses, and more |
| pymeshlab | lcc | 1.0 | PyMeshLab is a Python library that exposes the mesh-processing capabilities of MeshLab programmatically, letting users load, filter, analyze, and save 3D triangular meshes and point clouds from Python scripts. It centers on a MeshSet object to which meshes are added and to which the full catalog of MeshLab filters is applied by name with keyword parameters, covering cleaning, simplification/decimation, remeshing, surface reconstruction, curvature and geometric measurement, and format conversion across PLY, OBJ, STL, OFF and others. This makes it ideal for batch and reproducible pipelines that would be tedious in the GUI — for example loading a mesh, applying quadric edge-collapse decimation to a target face count, and saving the result. It is installed from PyPI in a Python 3.8 conda environment with NumPy for exchanging vertex/face arrays. | 3D modeling, Mesh processing, Computer graphics, Scientific computing | 3D Mesh Processing, Computer Science | Library | Documentation, Uses, and more |
| pypy | lcc | N/A | PyPy is a fast, compliant alternative implementation of the Python programming language. It aims to provide a high-performance, flexible, and efficient interpreter for Python, while maintaining compatibility with the CPython reference implementation. | Python Interpreter, Alternative Implementation, Just-In-Time Compilation, High Performance | Software Engineering, Computer & Information Sciences | Programming Language Implementation | Documentation, Uses, and more |
| pyradiomics | lcc | 1.0 | PyRadiomics is an open-source library for extracting quantitative imaging biomarkers, commonly called radiomic features, from medical images and their accompanying segmentation masks. From a paired image and region-of-interest label map it computes first-order intensity statistics, three-dimensional shape descriptors such as volume, surface area, sphericity and elongation, and a large family of texture measures derived from grey-level co-occurrence, run-length, size-zone, dependence and neighbouring-grey-tone-difference matrices. Features can additionally be recomputed on filtered versions of the image, including wavelet decompositions and Laplacian-of-Gaussian edge enhancements, which expands a single case to well over a thousand descriptors. Extraction is driven by a parameter file that records binning, resampling, interpolation and which feature classes are enabled, so that a study's settings are reproducible and reportable. It reads any format supported by SimpleITK, including the NRRD label maps produced by segmentation models, and can process many cases in one pass from a manifest. Feature extraction is processor-only and needs no GPU. | Documentation, Uses, and more | |||
| pysam | lcc | 1.0 | pysam (version 0.19.1 from bioconda) is a Python module that wraps the htslib C library (the engine behind samtools and bcftools), giving programmatic access to high-throughput sequencing data files from Python. It provides classes to read, write, and randomly access SAM/BAM/CRAM alignment files and VCF/BCF variant files, iterate over reads in a region (using the BAM index), inspect per-read fields (flags, CIGAR, tags, mapping quality), pile up columns for per-base analysis, and query FASTA/tabix-indexed files. It also exposes thin wrappers around samtools/bcftools subcommands. Rather than a command-line tool, it is a library used to build custom analyses and pipelines. It is one of the most heavily used building blocks in Python bioinformatics, underpinning many variant-calling, coverage, and QC scripts and tools. | Bioinformatics, Sequence Analysis, Sam/Bam Files, Python Module | Bioinformatics, Biological Sciences | Library | Documentation, Uses, and more |
| pyside6 | mcc | 1.0 | PySide6 is the official set of Python bindings for the Qt 6 application framework (the Qt for Python project), used to build cross-platform graphical desktop applications in Python. It exposes Qt's extensive C++ modules to Python — QtWidgets for conventional desktop UIs, QtCore/QtGui, Qt Quick/QML, charts, multimedia, and more — plus tooling to load .ui designer files and compile resources, enabling rich native-looking GUIs with signals/slots event handling. This particular build packages PySide6 6.11.1 in a Python 3.11 conda environment together with the scientific stack (numpy, scipy, matplotlib-base), requests and lxml, the pyqtdarktheme styling helper, and pycatima, and it critically bundles the full system xcb-util library stack that Qt's 'xcb' platform plugin needs to display windows under X11. It is intended for running Qt Python GUI applications inside Open OnDemand Interactive Desktop sessions on the cluster. | Documentation, Uses, and more | |||
| python | mcc, lcc, ecc | N/A | Python 3.10 installed via Conda environment. | Programming Language, Interpreted Language, High-Level Language | Software Engineering, Computer & Information Sciences, Systems & Development | Language | Documentation, Uses, and more |
| python-venv | ecc | N/A | A tool for creating isolated Python environments, allowing users to manage dependencies for different projects independently. | Python, Virtual Environment, Dependency Management | Computer Science, Software Engineering | Library | Documentation, Uses, and more |
| python312_ml | mcc | 1.0 | This is a curated Python 3.12 machine-learning and data-science environment assembled from conda-forge and PyTorch channels rather than a single tool, intended as a batteries-included workspace for tabular ML and deep learning. It bundles scikit-learn for classical modeling, the leading gradient-boosting frameworks XGBoost, LightGBM, and CatBoost, model-interpretability via SHAP, hyperparameter optimization via Optuna, imbalanced-learn for resampling skewed classification datasets, and PyTorch for neural networks, alongside pandas for data wrangling and seaborn/plotly for static and interactive visualization. It suits end-to-end workflows: exploratory analysis, feature engineering, cross-validated model training and tuning, and explainability reporting, all from Python scripts or notebooks. Because everything is pinned into one environment, users avoid dependency conflicts between the boosting libraries and can move directly from a pandas DataFrame to trained, tuned, and interpreted models. | Documentation, Uses, and more | |||
| pytorch | lcc, ecc | 1.0 | PyTorch is a widely used open-source deep learning framework for building, training, and deploying neural networks, provided here as the NVIDIA NGC PyTorch container tuned for GPU acceleration. The NGC build ships PyTorch with CUDA, cuDNN, NCCL, and mixed-precision (AMP) support preconfigured, along with the usual scientific Python companions, so models train efficiently on the ECC GPUs without any local environment setup. It supports the full research workflow: tensor computation, automatic differentiation, distributed and multi-GPU training, and export of trained models for inference. Typical inputs are datasets and training scripts; outputs are trained model checkpoints, metrics, and predictions. This is the baseline PyTorch environment on ECC and the recommended starting point for GPU deep learning work on the cluster. | Machine Learning, Deep Learning, Numerical Computation, Neural Networks | Artificial Intelligence & Intelligent Systems, Computer & Information Sciences | Machine Learning Framework | Documentation, Uses, and more |
| pytraj | mcc, lcc | 1.0 | pytraj is a Python front end to cpptraj, the trajectory-analysis engine of AmberTools, providing fast and memory-efficient analysis of molecular-dynamics simulation trajectories. It reads the many trajectory and topology formats cpptraj supports (Amber NetCDF/mdcrd, PDB, DCD, and more) and exposes a NumPy-friendly, interactive API so trajectories can be loaded, sliced, and analyzed in scripts or Jupyter notebooks. It computes a broad range of structural and dynamical quantities — RMSD and RMSF, radius of gyration, distances/angles/dihedrals, hydrogen bonds, radial distribution functions, secondary structure, clustering, principal-component/normal-mode analysis, and more — often with the underlying C++ speed and optional parallelism. This is version 2.0.5 from the ambermd conda channel. | Molecular Dynamics, Bioinformatics, Computational Chemistry, Python | Molecular Dynamics, Biochemistry and Molecular Biology | Library | Documentation, Uses, and more |
| pyvcf | mcc | 1.0 | PyVCF is a Python library for parsing, filtering, and writing VCF (Variant Call Format) files, giving programmatic access to genetic variant records. It reads a VCF through a vcf.Reader, exposing each record's chromosome, position, reference and alternate alleles, quality, filter status, INFO fields, and per-sample genotype calls as Python objects, and provides a vcf.Writer for emitting VCF, enabling custom scripting of variant filtering, annotation, and format conversion where a full toolkit like BCFtools is unnecessary. This environment is the 0.6.8.dev0 development build and adds Biopython (for sequence handling, format parsing, and access to biological databases) and pandas (for tabular manipulation and analysis of extracted variant data), making it a convenient scripting environment for ad hoc variant and sequence analysis. Note that upstream PyVCF is unmaintained and targets Python 2/early Python 3; typical use is importing the library, opening a reader on a VCF file, and iterating over records. It is provided as a conda environment in a Rocky 8 multi-environment bioinformatics container. | bioinformatics, genomics, data analysis, variant calling | Genomics, Genetic Variation | Library | Documentation, Uses, and more |
| qiime2 | mcc, lcc | 1.0 | QIIME 2 is a widely used, plugin-based microbiome bioinformatics platform for analyzing marker-gene (amplicon) sequencing data such as 16S rRNA, 18S, and ITS. This deployment is the 2023.9 amplicon distribution and adds the q2-perc-norm plugin for percentile normalization of case/control microbiome studies to reduce batch effects. QIIME 2 emphasizes reproducibility and provenance tracking: all data are wrapped in typed .qza (artifact) and .qzv (visualization) files that record the full history of commands used to generate them. A standard workflow imports demultiplexed reads, denoises with DADA2 or Deblur to produce amplicon sequence variants (ASVs) and a feature table, assigns taxonomy against a reference (e.g. SILVA or Greengenes) with q2-feature-classifier, builds a phylogeny, and computes alpha/beta diversity metrics and differential-abundance results, all driven through the qiime command-line interface. Visualizations are viewed interactively at view.qiime2.org. The q2-perc-norm addition lets users normalize relative-abundance tables to percentiles within control samples before cross-study comparison. | Microbiome, Bioinformatics, Sequencing | Bioinformatics | Analysis Software | Documentation, Uses, and more |
| qiime2-amplicon | lcc | 1.0 | QIIME 2 (Quantitative Insights Into Microbial Ecology 2) is a community-standard, plugin-based platform for microbiome bioinformatics, and the amplicon distribution (release 2024.10) targets marker-gene surveys such as 16S rRNA, 18S, and ITS amplicon sequencing. It provides a fully reproducible workflow in which every step is tracked as provenance, wrapping tools for demultiplexing, quality filtering and denoising (DADA2, Deblur), feature-table construction, taxonomic classification (naive-Bayes classifiers, BLAST/VSEARCH), phylogenetic tree building, and diversity analyses (alpha/beta diversity, UniFrac, ordination). Data are handled as typed .qza artifacts and .qzv visualizations that can be viewed interactively at view.qiime2.org. Users typically interact through the qiime command-line interface (with plugins such as DADA2 denoising and phylogenetic core-metrics diversity), the Python 3 API, or Galaxy. It is well suited to end-to-end community-composition and differential-abundance studies from raw reads to publication-ready statistics and figures. | bioinformatics, microbiome, amplicon sequencing, data analysis | Bioinformatics, Biological Sciences | Open-source | Documentation, Uses, and more |
| qiime2-shotgun | lcc | 1.0 | QIIME 2 (shotgun distribution, release 2024.2, built from the official environment YAML) is a decentralized, plugin-based microbiome bioinformatics platform, and this particular distribution is configured specifically for shotgun metagenomics rather than amplicon analysis. It provides plugins for quality control and host filtering of metagenomic reads, taxonomic profiling (e.g., via Kraken2/Bracken integrations), functional profiling, and read/contig-based classification, all within QIIME 2's reproducible framework in which every step is tracked with full provenance. Data flow through typed QIIME 2 Artifacts (.qza) and Visualizations (.qzv), so intermediate and final results carry a complete, auditable record of the commands and parameters that produced them. It is used from the command line via the q2cli interface or through the Python API, and results can be viewed interactively at view.qiime2.org. Inputs are typically imported FASTQ metagenomic reads plus sample metadata; outputs include taxonomic and functional feature tables and visualizations suitable for downstream statistical and diversity analysis. It gives shotgun-metagenomics users the same provenance and reproducibility guarantees as the amplicon QIIME 2 workflow. | Documentation, Uses, and more | |||
| qiime2-tiny | lcc | 1.0 | QIIME 2 is a widely used, plugin-based platform for reproducible microbiome and amplicon (marker-gene) bioinformatics; this is the 'tiny' distribution of the 2023.9 release, a minimal install providing the QIIME 2 framework and a reduced set of core plugins with a much smaller dependency footprint than the full amplicon/metagenome distributions. It centers on the concept of QIIME 2 Artifacts (.qza data and .qzv visualizations) that embed provenance so every analysis step is tracked and reproducible. Even in the tiny build, the core framework supports importing data, managing artifacts, and running provenance-tracked commands via the qiime CLI (including tools import and the general plugin/action invocation pattern), with the Python API and view/export tooling available. It is intended for lightweight, scriptable pipelines and CI/container use where the full QIIME 2 distribution's size and dependency weight are unnecessary, while remaining upgradeable/extendable with additional plugins as needed. | Documentation, Uses, and more | |||
| qiime2amplicon | lcc | 1.0 | Documentation, Uses, and more | ||||
| qiskit | lcc | 1.0 | Qiskit is IBM's open-source Python SDK for programming quantum computers and simulators, spanning circuit construction, transpilation/optimization, and execution across simulators and real quantum hardware. Users build quantum circuits from gates and measurements with the QuantumCircuit API, then run them on backends — locally on simulators or remotely on IBM Quantum devices via primitives (Sampler and Estimator). This environment adds qiskit-aer-gpu, the GPU-accelerated build of the Aer high-performance simulator, which speeds up statevector, density-matrix, and other simulation methods for larger circuits by offloading the heavy linear algebra to CUDA-capable GPUs. It is used for quantum-algorithm development and research (e.g. VQE, QAOA, quantum machine learning, and error-mitigation experiments) and for teaching. A typical flow builds a circuit, transpiles it for a backend, and runs it via a primitive, with the GPU simulator selected by configuring the Aer backend with a GPU device. It targets quantum-computing research and prototyping on HPC hardware. | Quantum Computing, Open Source, Simulation, IBM | Quantum Computing, Quantum Algorithms | Library | Documentation, Uses, and more |
| quantum-espresso | mcc, lcc | 1.0 | Quantum ESPRESSO is an integrated, open-source suite for first-principles electronic-structure calculations and materials modeling based on density-functional theory (DFT), plane-wave basis sets, and pseudopotentials (norm-conserving, ultrasoft, and PAW). Its core program pw.x performs self-consistent-field total-energy calculations, structural relaxations, and molecular dynamics, while a rich set of companion executables extends it: ph.x for density-functional perturbation theory (phonons, electron-phonon coupling), pp.x and projwfc.x for post-processing charge densities and projected DOS, bands.x for band structures, neb.x for reaction paths, and more. Calculations are driven by structured input files specifying the system, k-point grid, pseudopotentials, and convergence parameters, and it is commonly run in parallel over MPI for large systems. This build is conda-forge qe 7.5, CPU and MPI-enabled. | Quantum Mechanics, Materials Science, Computational Physics | Chemistry, Condensed Matter Physics | Package | Documentation, Uses, and more |
| quantumespresso | lcc | 1.0 | Quantum ESPRESSO is an integrated suite for first-principles electronic-structure and materials modeling based on density-functional theory, plane-wave basis sets, and pseudopotentials, centered on pw.x for self-consistent total-energy, geometry-optimization, and molecular-dynamics calculations and complemented by ph.x (phonons and electron-phonon coupling via DFPT), pp.x, projwfc.x, bands.x, and other post-processing executables. This particular build is version 7.4.1, compiled through Spack on Rocky 9 with the Intel oneAPI compiler toolchain and linked against Intel oneAPI MPI and the Intel MKL math library, with libxc for extended exchange-correlation functionals and both MPI and OpenMP parallelism enabled — a performance-oriented CPU build for large DFT jobs on the cluster. Calculations are driven by structured input files and run in parallel, with band, phonon, and post-processing steps following. It complements the conda-forge QE 7.5 build with an Intel-optimized alternative. | Quantum Mechanics, Materials Science, Computational Physics | Chemistry, Condensed Matter Physics | Package | Documentation, Uses, and more |
| quast | lcc | 1.0 | QUAST (QUality ASsessment Tool for genome assemblies), version 5.0.2, is the de facto standard for evaluating and comparing the outputs of de novo genome assemblers. Given one or more assembly FASTA files, it reports core contiguity metrics including number of contigs, total length, largest contig, N50/NG50, L50, and GC content, and can operate reference-free or, when a reference genome and/or gene annotation (GFF) are supplied, compute reference-based statistics such as genome fraction, duplication ratio, number of misassemblies, mismatches, and indels per 100 kb. It produces an interactive HTML report plus plain-text, TSV, and PDF summaries, along with cumulative-length and Nx plots and (via the bundled Icarus) a contig alignment viewer. Specialized modes exist for metagenomes (MetaQUAST) and large eukaryotic genomes. | Genome Assembly, Bioinformatics, Quality Assessment | Genomics, Biological Sciences | Stand-Alone Tool | Documentation, Uses, and more |
| quorum | lcc | 1.0 | QuorUM (Quality Optimized Reads of the University of Maryland), version 1.1.1, is an error corrector for Illumina short reads that improves base accuracy prior to genome assembly. It builds a k-mer frequency database from the reads, treats high-frequency k-mers as trusted, and then walks each read replacing likely sequencing errors (substitutions and small indels) with bases supported by the trusted k-mer set, while trimming or discarding reads that cannot be reliably corrected. Correcting reads up front reduces assembly-graph complexity and can improve contiguity and accuracy, particularly for the overlap-based assemblers with which it is bundled. It takes paired FASTQ input with configurable k-mer size and thread count (directly or via its wrapper script), producing corrected FASTA/FASTQ output. QuorUM is developed by the MaSuRCA team and is often used as a preprocessing step within or alongside the MaSuRCA assembler. | Documentation, Uses, and more | |||
| r | mcc, lcc | 1.0 | R | Statistical Computing, Data Analysis, Statistical Graphics | Computer Science | Statistical Software | Documentation, Uses, and more |
| r-base | lcc | N/A | R is a programming language and free software environment for statistical computing and graphics supported by the R Foundation for Statistical Computing. | Statistics, Data Science, Visualization, Open Source | Statistics, Statistical Computing | Programming Language | Documentation, Uses, and more |
| r-colorspace | mcc | N/A | Carries out mapping between assorted color spaces including RGB, HSV, HLS, CIEXYZ, CIELUV, HCL (polar CIELUV), CIELAB and polar CIELAB. Qualitative, sequential, and diverging color palettes based on HCL colors are provided along with corresponding ggplot2 color scales. Color palette choice is aided by an interactive app (with either a Tcl/Tk or a shiny GUI) and shiny apps with an HCL color picker and a color vision deficiency emulator. Plotting functions for displaying and assessing palettes include color swatches, visualizations of the HCL space, and trajectories in HCL and/or RGB spectrum. Color manipulation functions include: desaturation, lightening/darkening, mixing, and simulation of color vision deficiencies (deutanomaly, protanomaly, tritanomaly). Description Source: https://github.com/conda-forge/r-colorspace-feedstock/blob/main/recipe/meta.yaml | Color Models, Color Conversion, Data Visualization, R Package | Statistics, Other Mathematics | Package | Documentation, Uses, and more |
| r-fastbaps | lcc | 1.0 | fastbaps is an R package that performs fast hierarchical Bayesian clustering of aligned sequences into distinct populations, serving as a much faster approximation of the popular BAPS (Bayesian Analysis of Population Structure) approach. Given a multiple-sequence alignment (typically a core-genome or SNP alignment of bacterial isolates), it uses a Dirichlet-process prior and an optimized hierarchical algorithm to partition samples into genetically coherent clusters, and can operate at several complexity levels or optimize the prior via a Bayesian information criterion. Core usage in R is to import the alignment into a sparse representation, build the clustering, and extract the best-supported level assignments; results are readily overlaid on a phylogeny for context. Version 1.0.8 is widely used in bacterial population genomics to define lineages/sequence clusters far more quickly than MCMC-based tools. | Documentation, Uses, and more | |||
| r-peer | lcc | 1.0 | PEER (Probabilistic Estimation of Expression Residuals) is a statistical method, provided here as an R package, for inferring and removing hidden sources of variation (unobserved confounders such as batch effects and technical or environmental factors) from high-dimensional gene-expression datasets. It fits a Bayesian factor-analysis model that learns a set of latent factors explaining broad, coordinated variability across genes and individuals; regressing these factors out yields corrected expression residuals that substantially improve the power and calibration of downstream analyses, most notably expression quantitative trait loci (eQTL) mapping. Users supply an expression matrix (samples by genes), set the number of hidden factors to learn (and optionally known covariates), run the model to convergence, and extract the learned factors and residuals for use in association testing. It is a standard preprocessing step in large expression-genetics studies such as GTEx. This is bioconda r-peer version 1.3. | Documentation, Uses, and more | |||
| r-plm-1.6_5 | lcc | 1.0 | plm is an R package for econometric analysis of panel (longitudinal) data, where observations are indexed by both cross-sectional unit and time. Its core plm() function estimates the standard linear panel models, including pooled OLS, fixed-effects (within), random-effects (via Swamy-Arora, Amemiya, Wallace-Hussain, or Nerlove estimators), between, and first-difference specifications, as well as instrumental-variable and certain dynamic (GMM, via pgmm) models. It provides the diagnostic machinery panel work requires: the Hausman test (phtest) to choose between fixed and random effects, Breusch-Pagan LM tests for individual/time effects, F-tests for poolability, and serial-correlation and cross-sectional-dependence tests, along with robust and clustered covariance estimators. Data are supplied as a data frame declared with index columns (or a pdata.frame), with the model type and effect structure selected as arguments to the fitting call. Version 1.6-5 is packaged here as one of several independent conda R environments in a multi-app Miniconda container on Rocky 9. | Documentation, Uses, and more | |||
| r-satscan | lcc | N/A | R-SaTScan couples the SaTScan spatial/space-time scan statistic software with R (via the rsatscan package) and Java. It detects and evaluates statistically significant geographic clusters of events, used heavily in epidemiology, public health surveillance, and spatial disease-cluster detection. Delivered here as a Conda environment bundling R 3.5.1, the rsatscan library, and the SaTScan Java engine. | Documentation, Uses, and more | |||
| r-seurat | mcc | 1.0 | Seurat is a widely used R toolkit for single-cell genomics analysis, particularly single-cell and single-nucleus RNA-seq, and multimodal data such as CITE-seq and spatial transcriptomics. It provides an end-to-end workflow covering quality-control filtering, normalization (including SCTransform), highly-variable feature selection, dimensionality reduction (PCA, UMAP, t-SNE), graph-based clustering, and differential-expression / marker identification. A signature capability is its anchor-based integration for correcting batch effects and jointly analyzing datasets across conditions or modalities. Data is organized around the Seurat object built from a genes-by-cells count matrix (e.g. Cell Ranger output); a typical session creates the Seurat object, normalizes or runs SCTransform, finds variable features, scales the data, runs PCA, finds neighbors and clusters, and runs UMAP, followed by visualization with feature plots, dimensional-reduction plots, and violin plots. This is version 5.0.1. | bioinformatics, single-cell analysis, R package, data visualization | Bioinformatics, Biochemistry and Molecular Biology | Library | Documentation, Uses, and more |
| r-wgcna | lcc | 1.0 | WGCNA (Weighted Gene Co-expression Network Analysis) is an R package for constructing gene co-expression networks from high-dimensional expression data (typically RNA-seq or microarray) and identifying modules of co-expressed genes. It computes pairwise gene-gene correlations, raises them to a soft-thresholding power to build a weighted adjacency matrix that emphasizes strong connections, derives a topological-overlap measure, and applies hierarchical clustering to detect modules; it then summarizes each module by its eigengene and correlates modules with external sample traits to find biologically meaningful gene groups, along with measures of module membership and hub-gene identification. This environment augments WGCNA (v1.69) with a set of companion genomics R/Bioconductor packages for a fuller expression-analysis workflow: DESeq2 and edgeR for differential-expression normalization and testing, SAMR for significance analysis of microarrays, LDlinkR for querying linkage-disequilibrium data, and ggplot2 for visualization. Analyses are run interactively or via Rscript. It is widely used in systems-biology studies to relate expression modules to phenotypes. | R, Bioinformatics, Gene Expression, Network Analysis | Bioinformatics, Biological Sciences | Statistical Software | Documentation, Uses, and more |
| racon | lcc | N/A | Racon is consensus module for raw de novo DNA assembly of long uncorrected reads. | Bioinformatics, Sequence Analysis, Genomics | Genomics, Biological Sciences | Bioinformatics Tool | Documentation, Uses, and more |
| ragtag | mcc, lcc | 1.0 | RagTag (bioconda 2.1.0) is a toolkit for reference-guided improvement of genome assemblies, using alignments to one or more high-quality reference genomes to organize and repair a draft assembly without altering the underlying sequence content beyond documented joins. Its main functions are: scaffold, which orders and orients contigs into scaffolds based on their alignment to a reference (inserting gap runs of Ns); correct, which detects and breaks likely misassembled contigs; patch, which fills gaps and extends/joins sequences using another assembly; and merge, which reconciles scaffolding evidence from multiple sources (including Hi-C linkage). It relies on a fast whole-genome aligner (minimap2/nucmer) under the hood and outputs updated FASTA assemblies along with AGP files documenting every scaffolding decision so changes are traceable. It is a lightweight, commonly used step for reaching chromosome-scale assemblies when a related reference is available. | Assembly Improvement, Genome Alignment, Assembly Evaluation | Biological Sciences | Bioinformatics | Documentation, Uses, and more |
| raisd | mcc | 1.0 | RAiSD (Raised Accuracy in Sweep Detection) is a fast, lightweight tool for detecting positive selection and selective sweeps from population-genomic SNP data. It computes a composite statistic, the mu (μ) statistic, which jointly captures the three genomic signatures of a sweep — a local reduction in polymorphism, a shift in the site-frequency spectrum toward rare and high-frequency derived variants, and a specific pattern of linkage disequilibrium — scanning along the genome to flag candidate sweep regions. Input is a VCF (or ms-style) file of SNP data, optionally with sample/population specification, and output is per-window μ scores plus reports identifying the strongest sweep signals; it is designed to scale to whole-genome, many-sample datasets efficiently. This build is compiled from the RAiSD source archive (github.com/alachins/RAiSD) via install-RAiSD.sh. | Documentation, Uses, and more | |||
| randrproto | mcc, lcc | N/A | Documentation, Uses, and more | ||||
| rapidjson | mcc | N/A | RapidJSON is a fast JSON parser/generator for C++ with both SAX/DOM style API. Description Source: https://github.com/Tencent/rapidjson/ |
Json, C++, Parser, Generator, Sax, Dom, Memory Efficient | Software Engineering, Computer & Information Sciences | Library | Documentation, Uses, and more |
| raxml | mcc, lcc | 1.0 | RAxML (Randomized Axelerated Maximum Likelihood) is a widely cited program for maximum-likelihood phylogenetic tree inference from large nucleotide or amino-acid multiple sequence alignments. It performs ML tree searches under substitution models such as GTR (with GAMMA-distributed rate heterogeneity) for DNA and many empirical models for proteins, supports rapid and standard bootstrapping for branch support, partitioned analyses for multi-gene datasets, and combined ML-search-plus-bootstrap runs. It reads PHYLIP/FASTA alignments plus an optional partition file and writes best-tree, bootstrap, and info files, and offers multithreaded and vectorized (e.g. PTHREADS-SSE3) binaries. This is version 8.2.12 (the classic RAxML, predecessor to RAxML-NG). | Phylogenetics, Evolutionary Biology, Bioinformatics | Systematics & Population Biology, Biological Sciences | Bioinformatics Tool | Documentation, Uses, and more |
| raxml-ng | lcc | 1.0 | RAxML-NG is a from-scratch successor to the classic RAxML for maximum-likelihood phylogenetic tree inference, redesigned for better usability, accuracy, and scalability on large multi-gene alignments. It searches for the tree topology and branch lengths that maximize the likelihood under nucleotide or amino-acid substitution models (GTR, GAMMA rate heterogeneity, partitioned models, etc.), and provides standard and Transfer-Bootstrap support estimation, bootstopping convergence criteria, model checking, and restart/checkpointing for long runs. It reads FASTA/PHYLIP alignments plus an optional partition file, and can precompile the alignment into an efficient binary (.rba) for repeated analyses; outputs include the best-scoring ML tree, per-run log-likelihoods, and support values mapped onto the tree. It can perform an all-in-one analysis that combines the ML search with bootstrapping in a single step. Version 1.2.2 is provided as a conda tool within a Rocky 9 multi-tool bioinformatics container. | Phylogenetics, Evolutionary Biology, Bioinformatics | Systematics & Population Biology, Biological Sciences | Bioinformatics Tool | Documentation, Uses, and more |
| raxml-ng-mpi | mcc | N/A | raxml-ng software. | phylogenetics, maximum likelihood, bioinformatics, high-performance computing | Bioinformatics, Phylogenetics | Command-line tool | Documentation, Uses, and more |
| raxml-ng_v0.9.0 | lcc | N/A | RAxML-NG is a phylogenetic tree inference tool that performs maximum-likelihood analysis of large sequence alignments. A rewrite of the classic RAxML, it offers faster, more memory-efficient tree searches, bootstrap support estimation, and model selection. Used in molecular phylogenetics and evolutionary genomics for building species and gene trees. | Phylogenetics, Bioinformatics, Evolutionary Biology | Bioinformatics, Biological Sciences | Command-line tool | Documentation, Uses, and more |
| raxmlng | mcc | 1.0 | RAxML-NG, version 1.2.0, is a maximum-likelihood phylogenetic tree inference tool, a from-scratch redesign of the classic RAxML that is faster, more memory-efficient, and easier to use for very large sequence alignments. Built on the libpll phylogenetic likelihood library, it searches tree space for the ML topology under a chosen substitution model, supports a wide range of DNA and protein models (with automated model selection available via the companion ModelTest-NG), and provides bootstrapping with automatic convergence detection (MRE bootstopping) plus Transfer Bootstrap Expectation support values. It reads alignments in FASTA/PHYLIP (or a compressed binary format for speed) and a model specification, and offers an all-in-one mode that runs the ML search plus bootstrapping and maps support onto the best tree in a single step. Outputs include the best ML tree, bootstrap trees, and support-annotated trees in Newick format, making it a standard choice for large-scale ML phylogenetics. | phylogenetics, maximum likelihood, tree inference, bioinformatics | Bioinformatics, Biological Sciences | Executable | Documentation, Uses, and more |
| rclone | mcc | N/A | Rclone 1.74.4 - command-line tool for syncing files to and from cloud storage | Cloud Storage, File Synchronization, Command-Line Tool | Computer Science | Data Management | Documentation, Uses, and more |
| rcorrector | mcc | 1.0 | Rcorrector is a k-mer-based error-correction tool specialized for Illumina RNA-seq reads, where uneven, expression-dependent coverage makes generic genomic error correctors unsuitable. It builds a De Bruijn k-mer graph and uses a locally adaptive k-mer count threshold to distinguish genuine low-coverage transcripts from sequencing errors, correcting likely erroneous bases while preserving true low-abundance transcript signal. It processes paired or single FASTQ reads and outputs corrected FASTQ with reads tagged as corrected or flagged as unfixable. It is commonly used as a preprocessing step before de novo transcriptome assembly. This is version 1.0.4. | RNA-Seq, Error Correction, Bioinformatics, Transcriptomics | Bioinformatics, Biochemistry and Molecular Biology | Command-line tool | Documentation, Uses, and more |
| rdkit | mcc | 1.0 | RDKit is a comprehensive open-source cheminformatics toolkit, used here through its Python (and underlying C++) API for manipulating and analyzing molecular structures. It parses and writes SMILES, SMARTS, InChI, Mol/SDF and other formats; performs substructure searching, molecule standardization, and reaction handling; generates 2D depictions and 3D conformers (with force-field optimization via UFF/MMFF); and computes a large library of molecular descriptors and fingerprints (Morgan/ECFP, RDKit, MACCS, topological) for similarity searching and machine learning. It integrates tightly with pandas (bundled here) via PandasTools for cheminformatics dataframes, and is a foundational building block for QSAR modeling, virtual screening, and molecular property prediction pipelines. Typical use is scripting in Python, parsing a molecule from SMILES and then computing descriptors or fingerprints from it. Installed from conda-forge alongside pandas. | Cheminformatics, Molecular Modeling, Cheminformatics Toolkit | Chemical Sciences, Natural Sciences | Toolkit | Documentation, Uses, and more |
| rdma-core | mcc | N/A | rdma-core is a set of userspace libraries and tools for Remote Direct Memory Access (RDMA) and kernel-level support for RDMA networking. | RDMA, Networking, High Performance Computing, InfiniBand | Computer Science | Library | Documentation, Uses, and more |
| rdna_gap_cleaning.sh | mcc | 1.0 | Documentation, Uses, and more | ||||
| re2c | ecc | N/A | re2c is a free and open-source lexer generator for C/C++, Go and Rust with a focus on generating fast code. It compiles regular expression specifications to deterministic finite automata and encodes them in the form of conditional jumps in the target language. This approach is generally faster than table-based lexers, and the generated code is easier to debug and understand. Description Source: https://re2c.org/ |
Lexer, Scanner, Regular Expressions, Dfa | Software Engineering, Computer & Information Sciences | Tool | Documentation, Uses, and more |
| readline | mcc, lcc, ecc | N/A | The GNU Readline library provides a set of functions for use by applications that allow users to edit command lines as they are typed in. Both Emacs and vi editing modes are available. The Readline library includes additional functions to maintain a list of previously-entered command lines, to recall and perhaps reedit those lines, and perform csh-like history expansion on previous commands. | Gnu, Command Line Interface, Interactive Programs, Line Editing, History Manipulation | Software Engineering, Computer & Information Sciences | Command Line Tool | Documentation, Uses, and more |
| recon | mcc | 1.0 | RECON is a program for de novo identification and classification of repetitive element families directly from genomic sequence, without prior knowledge of the repeats present. It clusters and analyzes patterns of pairwise sequence similarity to delineate distinct repeat-element families and their boundaries, a foundational step in building custom repeat libraries. In practice it is rarely run standalone; it is a core engine within RepeatModeler, which iteratively invokes RECON (alongside RepeatScout, LTR discovery, and classification) to produce a species-specific consensus repeat library that RepeatMasker then uses to annotate and mask transposable elements. Here it is provided as part of the Dfam TE Tools (TETools) v1.9.5 environment, packaged from the upstream dfam/tetools:1.95 Docker image, which curates a consistent, ready-to-run transposable-element annotation stack (RepeatModeler, RepeatMasker, RECON, RepeatScout, and dependencies). It is used by genome-annotation projects to characterize repeat content before gene annotation. | Documentation, Uses, and more | |||
| recordproto | mcc | N/A | Documentation, Uses, and more | ||||
| reditools1 | mcc | 1.0 | REDItools is a suite of Python scripts for the detection and analysis of RNA-editing events — most commonly adenosine-to-inosine (A-to-I, read as A-to-G) editing — from RNA-seq data, optionally paired with matched DNA-seq to exclude genomic SNPs. It scans read pileups from alignments (BAM) against a reference, tabulating per-position base substitutions and applying quality, coverage, strand, and substitution filters to distinguish genuine editing sites from sequencing errors and polymorphisms. Its scripts (e.g. REDItoolDnaRna.py, REDItoolKnown.py, REDItoolDenovo.py) support de novo discovery, analysis at known editing positions, and combined DNA/RNA comparison, producing tab-delimited tables of candidate editing sites with frequencies and supporting counts. This build derives from the upstream claudiologiudice/rna_editing_protocol Docker image and adds bioconda htslib, samtools, pblat, and pysam to support the REDItools workflow. | Documentation, Uses, and more | |||
| reditools2 | mcc | 1.0 | REDItools2 is a Python tool for the systematic detection and quantification of RNA-editing events — such as A-to-I (read as A-to-G) and C-to-U edits — from RNA-seq data, redesigned from the original REDItools for speed and large-scale/HPC use. It scans coordinate-sorted BAM alignments position by position (using samtools/htslib via pysam) and, at each genomic site, tallies the reference and alternative base counts across reads, applying quality and coverage filters to distinguish genuine editing from sequencing errors or SNPs; the output is a per-position table reporting coverage, base distribution, edited fraction, and the inferred substitution type. Its key advance is a parallel implementation that can partition the genome into intervals and distribute the analysis across many cores or cluster nodes, making transcriptome-wide editing detection tractable on large datasets. It optionally uses a parallel wrapper and interval coverage files to drive the distributed analysis. It is used in RNA-editing and post-transcriptional-regulation research. | Documentation, Uses, and more | |||
| reindeer | lcc | 1.0 | REINDEER (REad Index for abuNDancE quERy) builds a compact, searchable index over a large collection of sequencing datasets and answers per-dataset abundance queries for k-mers and sequences. Unlike presence/absence indexes, REINDEER records k-mer counts, so a query sequence returns its estimated abundance in each indexed sample — enabling, for example, expression-level or coverage queries of a gene or transcript across thousands of RNA-seq or genomic read sets without re-aligning to each dataset. It works by indexing per-sample de Bruijn graphs and exploiting monotigs to store counts efficiently, keeping the index far smaller than the raw data. The workflow indexes a list of input read sets (typically preprocessed into k-mer count files) and then serves abundance queries for sequences supplied in FASTA. It is installed here from the kamimrcht GitHub repository. | Documentation, Uses, and more | |||
| renderproto | mcc, lcc | N/A | X Rendering Extension. This extension defines the protcol for a digital image composition as the foundation of a new rendering model within the X Window System. | X11, Rendering, Graphics Protocol, 2D Graphics, Compositing, Open Source, Unix | Library / Protocol Specification | Documentation, Uses, and more | |
| repeatafterme | mcc | 1.0 | RepeatAfterMe (the RAMExtend tool) supports the curation of transposable-element (TE) libraries by extending and refining TE consensus sequences. During manual repeat-library building, seed consensus models are often truncated and need to be extended into their flanking regions to capture the full length of an element; RAMExtend automates this by collecting genomic instances of a consensus, aligning their flanks, and building an extended consensus, iterating outward until the element boundaries are reached. This produces higher-quality, full-length TE consensus models for use with RepeatMasker and Dfam. It is part of the Dfam TE Tools (TETools) v1.9.5 environment and is located at /opt/RepeatAfterMe/RAMExtend within the container, which is packaged directly from the upstream dfam/tetools:1.95 Docker image alongside the rest of the TE annotation toolchain. | Documentation, Uses, and more | |||
| repeatmasker | mcc, lcc | 1.0 | RepeatMasker is a widely used program that screens DNA sequences for interspersed repeats and low-complexity regions — transposable elements, satellites, simple repeats, and the like — and soft- or hard-masks them so that downstream gene prediction and comparative analyses are not confounded by repetitive DNA. It searches query sequences against repeat libraries (Dfam HMM profiles and/or RepBase consensus sequences) using a configurable search engine (RMBlast, HMMER, or cross_match) and produces a detailed annotation of every repeat found, including family, divergence, and coordinates. Outputs include a masked FASTA, a tabular .out annotation file, and summary tables of repeat content by class. It is a cornerstone of genome-annotation pipelines. Here it is provided as part of the Dfam TE Tools (TETools) v1.9.5 curated transposable-element environment, packaged directly from the upstream dfam/tetools:1.95 Docker image, which bundles RepeatMasker with its companion tools and search engines. | Genomics, Bioinformatics, Sequence Analysis, Repeat Detection | Genetics, Biological Sciences | Bioinformatics Tool | Documentation, Uses, and more |
| repeatmodeler | mcc | 1.0 | RepeatModeler is a de novo transposable-element (TE) family discovery and modeling package that automatically identifies repeat families in a genome assembly and builds a consensus repeat library, without relying on prior annotation. It orchestrates several structural and repeat-finding engines (including RECON and RepeatScout, plus its LTR structural pipeline) to sample, align, and refine repeat family consensus sequences, and its output library is the standard input for masking with RepeatMasker. The typical workflow first builds a sequence database from the genome FASTA with the BuildDatabase utility and then runs RepeatModeler against that database, producing a classified families FASTA. This build is provided as part of the Dfam TE Tools (TETools) v1.9.5 curated environment, packaged directly from the upstream dfam/tetools:1.95 Docker image, which bundles RepeatModeler, RepeatMasker, and their dependencies together. | Genomics, Bioinformatics, Repeat Sequences, De Novo Identification | Bioinformatics, Biological Sciences | Research Tool | Documentation, Uses, and more |
| repeatscout | mcc | 1.0 | RepeatScout is a tool for the de novo discovery of interspersed repeat families in genomic DNA, building consensus sequences for repeats without relying on a pre-existing repeat library. It first tabulates high-frequency l-mers across the genome (via a build_lmer_table step) and then greedily extends seeds into consensus repeat family sequences, applying statistical criteria to decide when to stop extension, after which filtering scripts remove low-complexity and tandem repeats and elements with too few genomic copies. Its output is a FASTA library of repeat consensus sequences that is typically handed to RepeatMasker to annotate and mask repeats genome-wide. In this deployment RepeatScout is one component of the Dfam TE Tools (TETools) v1.9.5 environment, a curated transposable-element annotation stack that also bundles RepeatModeler, RepeatMasker, RMBlast, and related utilities, packaged directly from the upstream dfam/tetools:1.95 Docker image (the tool itself is provided by that image rather than separately installed in the def). | Genomics, Bioinformatics, Computational Biology | Biological Sciences | Genomic Sequence Analysis | Documentation, Uses, and more |
| revbayes | lcc | 1.0 | RevBayes (built from the revbayes repository, branch v1.4.0, with the TensorPhylo plugin) is a flexible framework for Bayesian phylogenetic inference and evolutionary modeling based on probabilistic graphical models. Rather than offering a fixed set of analyses, it exposes an interactive interpreted language, Rev, in which the user explicitly assembles the model as a graph of random variables, deterministic transformations, and data-clamped nodes, then runs MCMC over it — this makes it possible to specify highly customized models for phylogeny estimation, divergence-time dating, molecular and morphological evolution, biogeography, and diversification. The provided 'rb' binary comes in serial and MPI builds (with OpenMPI 4.1.8) so analyses can be parallelized across chains/replicates on HPC. This build is extended with the TensorPhylo plugin, which provides fast likelihood computation for state-dependent speciation-extinction (SSE-type) diversification models, enabling large or complex trait-dependent diversification analyses. Usage typically runs a Rev script that defines the model, moves, monitors, and MCMC; outputs are logged parameter and tree samples analyzed for convergence and summarized into annotated trees. | Phylogenetics, Bayesian Inference, Evolutionary Biology | Systematics & Population Biology, Biological Sciences | Probabilistic Programming Tool | Documentation, Uses, and more |
| revbayes-mpi | lcc | 1.0 | This is the MPI-parallel build of RevBayes 1.4.0 (the rb-mpi executable), a program for Bayesian phylogenetic and macroevolutionary inference using probabilistic graphical models. RevBayes replaces fixed "black box" analyses with its own interpreted language, Rev, in which the user explicitly specifies the model as a directed graphical model (defining nodes for tree topology, branch lengths, substitution models, clocks, and priors) and then runs MCMC to sample the posterior. This flexibility supports a wide range of analyses — substitution-model and partition inference, divergence-time (relaxed-clock) dating, diversification/birth-death models, biogeography, and discrete-trait evolution. The MPI build (compiled with OpenMPI 4.1.8) distributes the MCMC computation — for example running multiple chains or parallelizing likelihood evaluation — across processes for faster analysis of large datasets, and this build also bundles the TensorPhylo plugin for fast state-dependent diversification (SSE-type) models. It targets computationally intensive Bayesian phylogenetics on HPC clusters. | Documentation, Uses, and more | |||
| reviewer | mcc | 1.0 | REViewer (Read Extraction and Visualization of Sequencing Reads) is an Illumina tool for generating read-pileup visualizations that let a user manually inspect and validate short-tandem-repeat (STR) genotype calls produced by ExpansionHunter. Taking the BAM/CRAM of realigned reads that ExpansionHunter emits together with the variant catalog and the ExpansionHunter VCF, it phases the supporting reads and draws an SVG showing each read aligned across the repeat region and its flanks, making it visually clear how many reads support each allele's repeat count and whether the call is well supported or ambiguous. This is important for confirming pathogenic repeat expansions in disorders such as Huntington's disease and various ataxias. Version 0.2.7 is provided as a conda tool in a Rocky 9 baseline container of long-read assembly/curation and related genomics tools. | Documentation, Uses, and more | |||
| ribotin-hifiasm | mcc | 1.0 | Hifiasm integrated entry point of ribotin, which recovers ribosomal DNA (rDNA) arrays from a hifiasm assembly instead of a Verkko one. It reads the hifiasm unitig graph under a given assembly prefix, identifies the rDNA tangles there from k-mer matches to a reference rDNA together with graph topology, extracts the HiFi reads uniquely assigned to each tangle, and builds a per tangle consensus and reduced graph. Unlike the Verkko mode it needs the read files passed explicitly, and the same reads that were given to hifiasm should be used; adding ultralong ONT reads produces clustered morph consensuses with abundances as well. Tangle node lists may also be supplied manually for organisms where automatic tangle detection is unreliable. Each tangle is written to its own output folder alongside its graph, consensus and morph files. | Documentation, Uses, and more | |||
| ribotin-ref | mcc | 1.0 | Reference guided entry point of ribotin, an assembler for ribosomal DNA (rDNA) arrays that works directly from sequencing reads with no genome assembly required. It recruits rDNA reads out of a HiFi or duplex read set by k-mer matching against a set of reference rDNA sequences, builds a de Bruijn graph of the recruited reads, and reports the most covered cycle through that graph as the rDNA consensus together with the graph reduced around it. Supplying ultralong Oxford Nanopore reads in addition makes it error correct those reads against the graph, extract the individual rDNA units they span, cluster those by sequence similarity and produce a consensus and an abundance estimate for each morph, which is how the rDNA arrays of the CHM13 reference were resolved. A human preset supplies the CHM13 recruitment sequences, the k-mer size, the morph clustering thresholds and the KY962518.1 orientation reference and annotation, and annotations are lifted onto both the consensus and each morph; other species are handled by giving an example morph file and an approximate morph size instead. Outputs are the consensus sequence, the read and reduced graphs in GFA, read and consensus paths in GAF, the morph consensuses with their coverages, a graph of how the morphs connect, and the lifted annotations in GFF3. | Documentation, Uses, and more | |||
| ribotin-verkko | mcc | 1.0 | Verkko integrated entry point of ribotin, for recovering ribosomal DNA (rDNA) arrays from a finished telomere to telomere assembly rather than from raw reads. rDNA arrays normally collapse into unresolved tangles in an assembly graph, and this mode locates those tangles in a Verkko assembly directory by combining k-mer matches to a reference rDNA with assembly graph topology, then pulls out the HiFi reads uniquely assigned to each tangle and assembles each one separately into a consensus. It reuses the ultralong ONT reads that the Verkko run itself was given, so morph consensuses and their abundances come out automatically for assemblies built with ONT data. Tangles can also be supplied by hand, one node list per tangle, when the automatic search is not appropriate for the organism. Results are written to one output folder per tangle, making it the natural companion to Verkko for finishing the rDNA regions that whole genome assembly leaves unresolved. | Documentation, Uses, and more | |||
| rmblast | mcc | 1.0 | RMBlast is a customized fork of the NCBI BLAST+ engine created specifically to support RepeatMasker and RepeatModeler, the standard tools for identifying and masking transposable elements and other repetitive sequences in genomes. It augments the standard rmblastn/blastn search program with the complex-alignment scoring features these repeat tools require — notably support for asymmetric/complexity-adjusted scoring matrices and the cross-match-style output that RepeatMasker consumes — so that repeat libraries such as Dfam and RepBase can be searched sensitively against genomic sequence. RMBlast is not typically invoked directly by end users; instead RepeatMasker/RepeatModeler call rmblastn internally as their search engine. This entry is provided as part of the Dfam TE Tools (TETools) v1.9.5 environment, packaged from the upstream dfam/tetools:1.95 Docker image, which curates RMBlast together with RepeatMasker, RepeatModeler, and their dependencies into a ready-to-run transposable-element annotation toolkit. | bioinformatics, sequence alignment, BLAST, high-performance computing | Computational Biology, Biochemistry and Molecular Biology | Command-line tool | Documentation, Uses, and more |
| rnabloom | mcc | 1.0 | RNA-Bloom is a reference-free de novo transcriptome assembler that reconstructs full-length transcript sequences directly from RNA sequencing reads without a reference genome. It uses memory-efficient Bloom-filter-based de Bruijn graph representations to handle large datasets, and supports both short-read (Illumina) assembly and, notably, long-read (Oxford Nanopore and PacBio) assembly through its RNA-Bloom2 long-read mode, making it versatile across sequencing platforms. Inputs are FASTQ reads (single or paired, or long reads), and outputs are assembled transcript FASTA files representing the sample's expressed isoforms; short-read runs take paired left/right reads while long reads use a dedicated long-read mode. It is often used when studying non-model organisms lacking a good reference. This is bioconda rnabloom 2.0.1. | Documentation, Uses, and more | |||
| rnahybrid | mcc | 1.0 | RNAhybrid (version 2.1.2) predicts the most favorable hybridization sites between a short RNA (typically a microRNA) and a longer target RNA by finding the minimum free energy of duplex formation, treating the problem as an extension of RNA secondary-structure energy minimization restricted to intermolecular base pairing. It is a core tool for microRNA target prediction, reporting the optimal (and suboptimal) hybridization duplexes, their free energies, and target positions, and can compute statistical significance (p-values) of predicted sites based on an extreme-value distribution fit to length- and composition-matched background sequences. Options allow forcing a seed match, setting energy cutoffs, and disallowing G:U wobble or bulges in the seed region. Inputs are the target sequence(s) and the query miRNA(s) in FASTA, and it produces the predicted duplexes and energies. It is used in regulatory-RNA and gene-regulation studies to nominate candidate miRNA-target interactions. | bioinformatics, RNA, hybridization, computational biology | Bioinformatics, Molecular Biology | Command-line tool | Documentation, Uses, and more |
| rnammer | lcc | 1.0 | RNAmmer predicts ribosomal RNA (rRNA) genes in genomic DNA sequences using hidden Markov models, identifying the 5S/8S, 16S/18S, and 23S/28S rRNA subunits across the bacterial, archaeal, and eukaryotic kingdoms. It applies kingdom- and molecule-specific HMM profiles (built via HMMER) to scan input sequence and report the locations, orientations, and scores of predicted rRNA genes, and is a common component of prokaryotic genome annotation and used, for example, to supply rRNA predictions to tools like Trinotate. Input is genomic FASTA with the kingdom and rRNA molecule type specified; outputs include GFF annotations and optionally extracted rRNA FASTA sequences. RNAmmer is academic-licensed software and is installed here from a locally staged archive (rnammer-1.2-ccs.tar) because it cannot be redistributed through public package channels. | bioinformatics, genomics, RNA, annotation | Bioinformatics, Biological Sciences | Standalone | Documentation, Uses, and more |
| rnastructure | lcc | 1.0 | RNAstructure is a software package for predicting and analyzing RNA (and DNA) secondary structure and its thermodynamics, used in structural and functional RNA biology. Its core predicts the minimum-free-energy structure and suboptimal structures using nearest-neighbor thermodynamic parameters, and it can compute base-pair probabilities and maximum-expected-accuracy structures via the partition function, fold two sequences together (bimolecular/hybridization), predict target accessibility for oligo/siRNA design, and incorporate experimental restraints such as SHAPE chemical-probing data as folding constraints. It reads FASTA/sequence input and outputs CT/dot-bracket structure files and drawings, through command-line programs such as Fold (structure prediction), partition (partition-function calculation), ProbabilityPlot, and bifold (bimolecular folding), plus a GUI. This is version 6.1. | RNA, Bioinformatics, Computational Biology, Molecular Biology | Bioinformatics, Biochemistry and Molecular Biology | Standalone | Documentation, Uses, and more |
| roary | lcc | 1.0 | Roary is a high-speed, stand-alone pan-genome pipeline for bacteria that, given a set of annotated assemblies, computes the pan-genome — the full complement of genes across the samples — partitioned into core genes (present in nearly all isolates) and accessory genes (present in only some). It takes GFF3 files (such as those produced by Prokka) as input, clusters coding sequences using CD-HIT followed by all-against-all BLASTP and Markov clustering (MCL), and scales to thousands of genomes efficiently. Outputs include a gene presence/absence matrix, a core-gene alignment (optionally via MAFFT) suitable for phylogenetics, and summary statistics and plots of pan-genome structure. In this environment the conda env additionally provides Kraken2 (2.0.9beta) for taxonomic classification and R 4.0.2 with ggplot2 for downstream plotting. This is bioconda roary version 3.13.0. | Bioinformatics, Genomics, Pan-Genome Analysis, Prokaryotic Genomes | Genomics, Biological Sciences | Tool | Documentation, Uses, and more |
| root | mcc, lcc | 1.0 | ROOT is CERN's C++/Python data-analysis framework, the de facto standard for processing and analyzing the very large datasets produced in high-energy and nuclear physics, though it is also used in astronomy and other data-intensive fields. It provides an efficient columnar storage format (TTree/RNTuple and .root files) for petabyte-scale datasets, a rich set of histogramming and multidimensional binning classes, curve fitting and minimization (via Minuit), statistical and multivariate analysis tools, and extensive 1D/2D/3D graphics for visualization. It ships an interactive C++ interpreter (Cling) plus PyROOT bindings so analyses can be scripted in Python, and RDataFrame offers a modern, declarative, parallelizable analysis interface. Users typically read event data from ROOT files, fill and fit histograms, and produce plots either interactively or in batch macros. This is version 6.34.4 from conda-forge. | Data Analysis, Visualization, High-Energy Physics, Data Manipulation, Statistical Analysis | Particle & High-Energy Physics, Physical Sciences | Library | Documentation, Uses, and more |
| rosetta | lcc | N/A | rosetta software. | bioinformatics, computational biology, structural biology, molecular modeling | Bioinformatics, Structural Biology | Application | Documentation, Uses, and more |
| rsem | mcc, lcc | 1.0 | RSEM (RNA-Seq by Expectation-Maximization, wrapped here from the dceoy/rsem Docker image) quantifies gene- and isoform-level expression from RNA-seq reads by using a statistical model that probabilistically assigns multi-mapping reads across transcript isoforms via the EM algorithm, rather than discarding them. It works against a transcriptome (with or without a reference genome, and can be run reference-free with a de novo assembly), first preparing references with its rsem-prepare-reference step (optionally invoking an aligner like Bowtie/Bowtie2/STAR) and then estimating abundances with its rsem-calculate-expression step from FASTQ reads or a transcriptome-coordinate BAM. Outputs include per-gene and per-isoform results with expected counts plus TPM and FPKM values, along with credibility intervals when run in Bayesian mode. These outputs feed directly into differential-expression tools such as DESeq2/edgeR. It has long been a standard, well-validated quantifier for isoform-aware expression analysis. | RNA-Seq, Transcriptomics, Gene Expression, Bioinformatics | Genomics, Biological Sciences | Computational Tool | Documentation, Uses, and more |
| rstan | lcc | N/A | R | Bayesian Statistics, MCMC, Statistical Modeling, Probabilistic Programming | Statistical Inference, Bayesian Statistics | Statistical Software | Documentation, Uses, and more |
| rstudio | mcc, lcc | 1.0 | RStudio Server is a browser-based integrated development environment for the R statistical programming language, provided here via the rocker/rstudio Docker base image alongside R 4.5.1. It delivers the familiar RStudio interface — source editor with syntax highlighting and execution, interactive console, environment/history panes, integrated plotting, package management, and help — accessed through a web browser rather than a desktop application, which makes it well suited to HPC and containerized deployments (e.g., launched as an Open OnDemand interactive session). Users write and run R scripts, R Markdown, and notebooks, manage workspaces, and visualize results without local installation. This container image comes preloaded with common statistical and machine-learning R packages so analyses can begin immediately. The underlying environment is standard R 4.5.1, so any CRAN/Bioconductor package can be installed as needed. | Ide, Statistical Computing, Data Visualization, Programming | Statistics, Other Natural Sciences | Ide | Documentation, Uses, and more |
| rstudio-server | lcc | 1.0 | RStudio-Server is a catalog stub registering the browser-based RStudio Server interactive application (app id rstudio-ood-lcc) offered through the LCC cluster's Open OnDemand (OOD) portal, rather than a container image. It launches a full RStudio IDE session on an LCC compute node, giving researchers the familiar RStudio environment — script editor, console, plots, environment/history panes, and package management — directly in a web browser without any local installation or X-forwarding. OOD handles submitting the underlying SLURM job, provisioning the requested CPUs, memory, GPUs, and walltime, and securely proxying the RStudio web interface back to the user. Through the OOD form the user selects the R version/module, resource requests, and partition, then opens the session and works interactively against cluster storage and compute. It is the standard way to do interactive R-based data analysis, statistics, and visualization on LCC. This registration exists so the app is discoverable in the SDS catalog. | Documentation, Uses, and more | |||
| rsync | mcc | N/A | rsync is a fast and versatile file copying tool that can synchronize files and directories between two locations over a network or locally. | file transfer, synchronization, backup, networking | Computer Science | Command-line tool | Documentation, Uses, and more |
| ruby | mcc | N/A | Ruby is a dynamic, open source programming language with a focus on simplicity and productivity. It has an elegant syntax that is natural to read and easy to write. \r Description Source: https://www.ruby-lang.org/en/ |
Programming Language, Scripting Language, Web Development | Computer Science | Programming | Documentation, Uses, and more |
| rust | mcc | N/A | Rust is a systems programming language that is designed to be fast, memory-efficient, and safe. It is often used for building high-performance applications, such as operating systems, web browsers, game engines, and embedded systems. | Programming Language, System Programming, Concurrency | Software Engineering, Computer & Information Sciences | Compiler | Documentation, Uses, and more |
| rvtests | mcc | 1.0 | rvtests (Rare Variant tests) is a command-line software package for genetic association analysis of sequencing data, supporting both common single-variant tests and, in particular, gene- or region-based tests aggregating rare variants. It efficiently reads genotypes from VCF/BCF (and supports genotype dosages/imputed data) together with phenotype and covariate files, and can perform single-variant score/Wald/likelihood-ratio tests as well as a comprehensive suite of rare-variant aggregation tests including burden tests (CMC, Zeggini, Madsen-Browning), variance-component tests (SKAT, SKAT-O), and others, using a supplied grouping (gene/set) definition. It also supports meta-analysis by generating the score statistics and covariance matrices consumed by tools like RareMETALS/rareMETAL. It takes an input VCF along with phenotype, covariate, and gene-definition files and writes association results for the chosen tests. It is used in statistical-genetics studies to test the aggregate effect of rare variants within genes. This deployment builds rvtests version 2.1.0 from the source tarball on Rocky 9. | genetics, bioinformatics, statistical analysis, rare variants | Bioinformatics, Genetics | Analysis Tool | Documentation, Uses, and more |
| sabre | lcc | 1.0 | Sabre is a lightweight barcode demultiplexing tool that separates a multiplexed FASTQ file (or pair of files) into per-sample FASTQ files based on 5' barcode sequences. It reads a barcode data file mapping each barcode to output filename(s), compares the start of each read against those barcodes allowing a configurable number of mismatches, trims the matched barcode from the read, and writes the read to the corresponding sample file while collecting reads that match no barcode into an 'unknown' file. It has single-end and paired-end modes; in paired-end mode the barcode is expected on the first read and both mates are routed together, with the barcode file listing each barcode and its output file. Version 1.000 is provided as a conda tool in a CentOS 8 baseline bioinformatics/cheminformatics container. | Documentation, Uses, and more | |||
| salmon | mcc, lcc | 1.0 | Salmon, version 1.6.0, performs fast and bias-aware quantification of transcript-level expression from RNA-seq reads without requiring a traditional base-to-base genome alignment. It uses selective alignment / quasi-mapping against a transcriptome index and a dual-phase inference algorithm to estimate transcript abundances, while explicitly modeling and correcting for GC content, sequence-specific, and positional fragment biases to improve accuracy. It supports two modes: mapping-based quantification directly from FASTQ reads, in which an index is built from a transcriptome and reads are then quantified against it, and alignment-based quantification from a transcriptome BAM. Its recommended decoy-aware index (including the genome as decoy sequence) reduces spurious mapping. Output is per-transcript estimated counts and TPM (quant.sf), which are commonly imported via tximport for gene-level differential-expression analysis in DESeq2/edgeR. It is a standard, high-throughput tool for bulk RNA-seq quantification. | RNA-Seq, Transcript Expression Quantification | Genetics, Biological Sciences | Tool | Documentation, Uses, and more |
| salome | mcc, lcc | N/A | salome software. | Open Source, Simulation, Pre-processing, Post-processing, Mesh Generation | Numerical Simulation, Engineering and Technology | Framework | Documentation, Uses, and more |
| salsa2 | mcc | 1.0 | SALSA2 (bioconda 2.3) is a genome-scaffolding tool that uses Hi-C chromatin-contact data to order and orient contigs into chromosome-scale scaffolds, while also correcting input misassemblies. It leverages the property that Hi-C contact frequency decays with genomic distance to link and orient contigs, and importantly it uses the contact signal to detect and break likely misjoins in the input assembly before scaffolding, reducing propagation of upstream errors. It can incorporate the assembly graph (GFA) from the assembler to further constrain and validate scaffolding decisions. Inputs are the contig FASTA, its index, and Hi-C read alignments provided as a BED file of alignments (derived from a name-sorted BAM), optionally with the assembly graph and a restriction-enzyme specification; outputs are scaffolded FASTA plus an AGP describing the scaffolds. It is commonly paired with hifiasm/purge_dups contigs to reach chromosome-level assemblies. | Documentation, Uses, and more | |||
| sam2 | lcc | 1.0 | Segment Anything Model 2 is the successor to SAM, built on a hierarchical image encoder and extended with a streaming memory design that was introduced for video, and it generally produces cleaner boundaries and better separation of adjacent structures than the original model. For still images it exposes the same working pattern as SAM: automatic generation of all plausible masks, or prompted segmentation of one structure, with results returned in an output format identical to SAM so that downstream analysis code needs no change when switching between them. Four checkpoint sizes are available, trading segmentation quality against speed, each paired with its own model configuration file. In this container it is the primary segmentation engine for three-dimensional bone structures in CT data, and its output feeds directly into radiomic feature extraction. Because the underlying PyTorch build targets newer GPU architectures, this model requires an A100 or H200 node. | Documentation, Uses, and more | |||
| sambamba | mcc | 1.0 | sambamba (version 0.8.2 from bioconda) is a high-performance toolkit for processing SAM/BAM/CRAM alignment files, written in D and multithreaded to exploit multiple cores for substantial speedups over single-threaded equivalents. Its subcommands cover the common alignment-manipulation tasks: view (filter/convert, with a powerful filter-expression language), sort (coordinate or name sorting), index, markdup (duplicate marking/removal), flagstat, and depth (per-base or per-region coverage). It is frequently used as a faster drop-in for samtools in variant-calling and coverage pipelines, particularly for sorting and duplicate marking of large BAMs. Its expressive filtering (e.g., by MAPQ, flags, tags) makes it convenient for extracting specific read subsets. In this container it sits alongside transcriptomics, single-cell, and docking tools as a general alignment utility. | Sam/Bam Files, High-Performance Tool, Parallel Processing, Genomics | Bioinformatics, Biological Sciences | Computational Software | Documentation, Uses, and more |
| samblaster | mcc | 1.0 | samblaster is a fast utility that streams a read-id-grouped SAM file (typically piped directly from an aligner such as BWA-MEM) and, in a single pass, marks or removes duplicate read pairs while simultaneously extracting discordant read pairs and split-read (chimeric/supplementary) alignments into separate SAM files. Those discordant and split-read outputs are exactly the signals downstream structural-variant callers (e.g. LUMPY) use to detect deletions, duplications, inversions, and translocations, so samblaster is a common preprocessing step in SV pipelines, and its duplicate marking is much faster and lower-memory than sort-based approaches because it exploits the fact that mates are adjacent in aligner output. Because it reads unsorted, name-grouped input, it is typically inserted inline in an alignment-to-sorting stream. Version 0.1.26 is provided as its own conda environment/app in a Miniconda base container bundling transcriptomics, single-cell, methylation, docking, and data-transfer tools. | Bioinformatics, Hpc Tools | Bioinformatics, Biological Sciences | Analysis Tool | Documentation, Uses, and more |
| samtools | mcc, lcc | 1.0 | SAMtools is a foundational suite of utilities for reading, writing, and manipulating high-throughput sequencing alignments in the SAM, BAM, and CRAM formats, and it is a core dependency of most variant-calling and alignment pipelines. It handles format conversion and compression (view), coordinate sorting (sort), indexing for random access (index), duplicate handling, merging (merge), alignment statistics (flagstat, stats, idxstats), read extraction by region, per-base depth (depth), and consensus/pileup generation (mpileup) that feeds downstream callers. It also indexes reference FASTA files (faidx). A canonical pipeline converts a SAM to BAM, sorts by coordinate, and indexes the result for random access. This is version 1.20, built on HTSlib. | Sequence Analysis, Genomics, Bioinformatics | Bioinformatics, Biological Sciences | Command Line Tool | Documentation, Uses, and more |
| samtools-bcftools-htslib | mcc | 1.0 | This is a combined environment bundling the three core components of the HTSlib ecosystem at matching version 1.21. SAMtools manipulates high-throughput sequencing alignments in SAM/BAM/CRAM format, providing sorting, indexing, merging, duplicate marking, filtering by flag/region/quality, statistics (flagstat, stats, depth, coverage), pileup generation, and format conversion. BCFtools operates on variant data in VCF/BCF format, performing variant calling from pileups (via the mpileup+call workflow), filtering, normalization (left-alignment and splitting of multiallelics), annotation, merging, concatenation, consensus generation, and per-sample statistics. HTSlib is the underlying C library both build on, supplying the compression and indexing (bgzip/tabix) that make coordinate-based random access to these files efficient. Keeping all three at 1.21 avoids version-mismatch issues in pipelines that chain them, for example sorting alignments with SAMtools and then calling variants by piping a BCFtools mpileup into a BCFtools call. They are provided as conda environments in a Rocky 9 multi-tool bioinformatics container. | bioinformatics, genomics, data analysis, sequence alignment | Bioinformatics, Biological Sciences | Command-line tool | Documentation, Uses, and more |
| sas | mcc, lcc | N/A | SAS | Documentation, Uses, and more | |||
| sbert | lcc | 1.0 | sentence-transformers (SBERT) is a Python framework for computing dense vector embeddings of sentences, paragraphs, and images using pretrained and fine-tunable transformer models (BERT, RoBERTa, MPNet, MiniLM, and multilingual variants, among others). It maps text into a fixed-dimensional semantic space where cosine similarity reflects meaning, enabling semantic search, clustering, paraphrase mining, retrieval-augmented generation, and duplicate detection at scale. The core API is straightforward: instantiate a pretrained model, encode a list of texts to obtain NumPy/torch embeddings, and compare them with cosine-similarity or semantic-search utilities. It also supports training/fine-tuning bi-encoders and cross-encoders with contrastive and triplet losses, and this build (5.4.1) ships with GPU-enabled PyTorch for fast batch encoding of large corpora. | Documentation, Uses, and more | |||
| scalapack | lcc | N/A | A subset of LAPACK routines redesigned for heterogenous computing | Linear-Algebra, Numerical-Computing, Distributed-Computing, Parallel-Computing | Other, Physical Sciences | Library | Documentation, Uses, and more |
| scalasca | mcc, lcc | N/A | Toolset for performance analysis of large-scale parallel applications | Performance Analysis, Parallel Applications, Performance Optimization, Mpi Applications, Visualization | Software Engineering, Systems & Development, Computer & Information Sciences | Performance Analysis Tool | Documentation, Uses, and more |
| scflow | lcc | N/A | scFLOW is a commercial Computational Fluid Dynamics (CFD) package from Software Cradle (Cradle CFD, now part of Cadence/Hexagon-MSC), NOT the single-cell RNA-seq R toolkit as the ambiguous name might suggest. It is a next-generation unstructured-mesh thermal-fluid solver used for aerospace/automotive aerodynamics, rotating equipment (fans/pumps), electronics cooling, multiphase and free-surface (VOF) phenomena, and heat transfer. The CCS install (ccs/SCFLow/2020) ships the Cradle toolchain: scflowpre (mesher/preprocessor), scflowsol (solver), scflowcomb, and scmonitor solver-monitor utilities, alongside CRADLE install/license configuration. | Documentation, Uses, and more | |||
| schrodinger | lcc | 1.0 | The Schrodinger Suites is a comprehensive commercial software platform for computational chemistry, molecular modeling, and structure-based drug discovery. It bundles a broad set of applications including Glide (protein-ligand docking and virtual screening at multiple precision levels HTVS/SP/XP), Jaguar (ab initio quantum-mechanical electronic-structure calculations), Prime (protein structure prediction, loop refinement, and MM-GBSA binding-energy estimation), Desmond (GPU-accelerated classical molecular dynamics), and FEP+ (relative binding free-energy perturbation for lead optimization), typically driven through the Maestro graphical interface or command-line job launchers. In this deployment the licensed 2025-2 installation itself is not baked into the image; it is bind-mounted at runtime from /opt/ohpc/pub/libs/ccs/Schrodinger_Suites_2025-2, and the container's job is to supply a modern glibc and the X11/runtime libraries needed so that this software runs correctly on LCC's older CentOS 7 compute nodes (both CPU and GPU). Use requires a valid Schrodinger license; jobs are usually submitted with the suite's own utilities and job-control launchers. | Molecular Simulations, Quantum Chemistry, Drug Discovery, Materials Research | Chemistry, Chemical Sciences | Commercial | Documentation, Uses, and more |
| scipy | lcc | N/A | SciPy is a collection of mathematical algorithms and convenience functions built on NumPy . It adds significant power to Python by providing the user with high-level commands and classes for manipulating and visualizing data. Description Source: https://docs.scipy.org/doc/scipy-1.8.0/tutorial/general.html |
Scientific Computing, Numerical Algorithms, Python Library, Data Analysis, Signal Processing | Applied Mathematics, Computer & Information Sciences, Other Computer & Information Sciences | Python | Documentation, Uses, and more |
| scons | mcc, lcc | 1.0 | SCons is a software construction (build automation) tool implemented in Python, intended as a modern replacement for make and autotools. Build configuration files (SConstruct/SConscript) are written in ordinary Python, so build logic uses a full programming language rather than a specialized makefile syntax. SCons determines what to rebuild using file content signatures (MD5) rather than timestamps, which makes incremental builds reliable, and it automatically scans C/C++ (and other) source files for header dependencies so dependency lists need not be maintained by hand. It has built-in support for many languages and toolchains, parallel builds, and cross-platform operation. Developers invoke it in a directory containing an SConstruct file. In this container it is present primarily as a build dependency for other bundled tools. This is version 3.1.2 from conda-forge. | Build Tool, Software Construction, Automation, Python | Software Development, Engineering & Technology | Build Tool | Documentation, Uses, and more |
| scorep | mcc, lcc | N/A | Scalable Performance Measurement Infrastructure for Parallel Codes | Performance Analysis, Profiling, Tracing, HPC | Computer Science | Performance Analysis Tool | Documentation, Uses, and more |
| scotch | mcc, lcc | N/A | Graph, mesh and hypergraph partitioning library | Graph Partitioning, Mesh Partitioning, Parallel Computing, Distributed Memory | Software Engineering, Systems & Development, Computer & Information Sciences | Library | Documentation, Uses, and more |
| screen | mcc | N/A | Screen is a terminal multiplexer that allows users to use multiple terminal sessions within a single window. It is particularly useful for managing long-running processes and for remote sessions. | Terminal, Multiplexer, Remote Access, Unix | Systems, and Development, Other Computer and Information Sciences | Utility | Documentation, Uses, and more |
| screen-assembly.sh | mcc | 1.0 | Documentation, Uses, and more | ||||
| scte | mcc | 1.0 | scTE quantifies transposable-element (TE) expression at single-cell resolution from single-cell RNA-seq data, filling a gap left by standard gene-level quantification pipelines that typically discard or misassign multi-mapping repeat-derived reads. It builds an index combining gene annotations with transposable-element annotations and then assigns aligned reads (from BAM files produced by 10x Cell Ranger, STARsolo, or similar) to genes and TE families, producing a cell-by-feature count matrix that includes TE loci/subfamilies alongside genes. This enables study of TE activity in individual cells and across cell types — relevant to development, aging, and disease. The workflow is two conceptual steps: a build step to construct the genome/TE index, then a counting step to quantify from a BAM and emit a matrix suitable for downstream single-cell tools. It uses pysam for BAM handling. This is version 1.0.0 from bioconda. | Documentation, Uses, and more | |||
| sda | mcc, lcc | 1.0 | SDA (Segmental Duplication Assembler, built from the mrvollger/SDA repository with a bioconda RepeatMasker and mamba dependency stack) is a specialized pipeline for resolving and assembling segmental duplications (SDs) and other high-identity, collapsed repetitive regions that standard assemblers fold into a single copy. It works from long-read data (originally PacBio) by identifying paralog-specific variants (PSVs) — positions that distinguish near-identical duplicate copies — and using a correlation-clustering approach over the reads carrying those variants to separate and locally assemble the individual paralogs. This recovers the true copy number and sequence of duplicated regions that are otherwise misrepresented in a collapsed assembly. It is implemented as a Snakemake workflow that orchestrates read alignment, PSV detection, graph-based clustering, and local assembly, with RepeatMasker used for repeat annotation. Inputs are long reads aligned to a reference/assembly and the collapsed region definitions; outputs are the resolved paralogous contigs. It is used in structural- and comparative-genomics studies of duplication-rich genomic regions. | Documentation, Uses, and more | |||
| sed | ecc | N/A | sed (Stream Editor) is a Unix utility that parses and transforms text from a data stream or file using a simple, compact programming language. | Text processing, Unix, Command line, Scripting | Software Engineering, Other Computer and Information Sciences | Command-line tool | Documentation, Uses, and more |
| segment-anything | lcc | 1.0 | Segment Anything (SAM) is Meta AI's promptable image segmentation model, trained on a very large mask corpus so that it generalises to objects and imaging modalities it was never explicitly trained on. Given a single image it can either segment everything automatically, producing a set of candidate masks ranked by predicted quality and stability, or segment a specific structure in response to a prompt such as a point, a box or a coarse mask. Each returned mask carries its pixel area, bounding box, predicted intersection-over-union and a stability score, which makes it straightforward to filter or rank candidates downstream. In this container it is used for slice-by-slice segmentation of computed-tomography bone imagery, where the resulting per-slice masks are stacked into three-dimensional label maps. The ViT-H checkpoint is stored outside the image and referenced by absolute path, and the model has been patched for the stricter checkpoint-loading behaviour introduced in recent PyTorch releases. | Documentation, Uses, and more | |||
| seqkit | mcc | 1.0 | SeqKit is a fast, cross-platform, ultra-versatile command-line toolkit for manipulating FASTA and FASTQ sequence files, written in Go for high performance and streaming operation. It offers a large collection of subcommands for everyday sequence wrangling: stats (summary statistics), seq (transform/filter/reverse-complement), subseq (extract regions), grep (search by name/pattern/motif), fx2tab/tab2fx (convert to and from tabular form), split/split2 (partition files), sample and head (subsampling), rmdup (remove duplicates), sort, replace, and translate, among many others. It reads and writes gzip transparently, handles both file input and stdin/stdout for easy pipelining, and works comfortably on very large datasets. This is version 2.9.0 from bioconda. | Sequence Manipulation, Fasta, Fastq | Bioinformatics, Biological Sciences | Bioinformatics | Documentation, Uses, and more |
| seqsero2 | lcc | 1.0 | SeqSero2 is a bioinformatics tool for in silico serotyping of Salmonella enterica, predicting serotypes directly from whole-genome sequencing data as a rapid replacement for traditional antisera-based (Kauffmann-White) typing. It determines the O and H antigen formulas and maps them to a serotype name, and it can operate from raw Illumina reads, other read types, or assembled genomes. It offers both a k-mer-based mode for speed and a microassembly/allele mapping mode for higher accuracy, with a data-type selector distinguishing single/paired reads or assembly input. It produces a report with the predicted antigenic profile and serotype. This is version 1.3.1. | Documentation, Uses, and more | |||
| seqtk | mcc | 1.0 | seqtk is a fast, lightweight C toolkit by Heng Li for common manipulations of FASTA and FASTQ sequence files, widely used as a Unix-style building block in genomics pipelines. Invoked as subcommands, it handles tasks such as random subsampling of reads with a fixed seed (the sample subcommand), converting FASTQ to FASTA (the seq subcommand), trimming low-quality ends or fixed lengths, reverse-complementing, extracting subsequences by name or region, masking, and computing basic composition statistics (the comp subcommand). It streams efficiently and handles gzipped input, making operations on large read sets quick and memory-light; for example the sample subcommand can downsample to a fixed number of reads reproducibly, and the seq subcommand can convert between formats. It is a near-ubiquitous helper for preprocessing and reformatting sequence data. This is bioconda seqtk 1.4. | bioinformatics, sequence processing, FASTA, FASTQ, data manipulation | Bioinformatics, Biological Sciences | Command-line tool | Documentation, Uses, and more |
| shapeit2 | mcc | 1.0 | SHAPEIT2 (Segmented HAPlotype Estimation and Imputation Tool, release v2.r904) estimates the haplotype phase of individuals from unphased genotype data — that is, it resolves which alleles at nearby heterozygous sites lie on the same parental chromosome. It uses a fast, accurate model based on conditioning each individual's haplotypes on a set of others (a hidden-Markov/segmented approach) and can incorporate a recombination map, and it scales to genome-wide SNP-array or sequencing genotype data across large cohorts. Its primary role is as the pre-phasing step ahead of genotype imputation (feeding tools like IMPUTE2 or Minimac), and it can also phase using reference panels or family/trio information. Inputs are genotypes in PLINK (PED/MAP or BED) or GEN/sample format plus a genetic map; the output is phased haplotypes (.haps/.sample). This build is a precompiled Linux binary (glibc 2.17). It remains widely used in statistical and population genetics. | Genetics, Bioinformatics, Phasing, Genotype | Bioinformatics, Genetics | Command-line tool | Documentation, Uses, and more |
| shapeit5 | mcc | 1.0 | SHAPEIT5 is a statistical phasing tool that estimates haplotypes from genotype data in very large cohorts, with a particular emphasis on accurately phasing rare variants. It splits the phasing problem into a common-variant step (phase_common, using a Positional Burrows-Wheeler Transform / Hidden Markov Model approach) and a dedicated rare-variant step (phase_rare) that phases scarce alleles onto an already-phased common-variant scaffold, plus a ligate step to stitch phased chunks together. This design scales to hundreds of thousands of samples (e.g., UK Biobank-scale WGS/WES) while retaining accuracy at low allele frequencies. Inputs are typically genotypes in VCF/BCF format (optionally with a genetic map and reference panel) and outputs are phased haplotypes in BCF/VCF, which then feed imputation or downstream haplotype-based analyses. This is version 5.1.1 from bioconda. | Genomics, Data Phasing, Haplotype Inference | Genetics, Biological Sciences | Phasing Tool | Documentation, Uses, and more |
| shared-mime-info | mcc | N/A | A database of shared MIME types used to identify file types based on their content and extensions. | MIME types, file identification, desktop environments, file extensions | Software Engineering, Other Computer and Information Sciences | Library | Documentation, Uses, and more |
| shasta | lcc | N/A | shasta software. | Genome Assembly, Long-Read Sequencing, Bioinformatics | Genetics, Biological Sciences | Bioinformatics Tool | Documentation, Uses, and more |
| shtools | lcc | N/A | SHTOOLS is a software package for the analysis and synthesis of spherical harmonic functions, widely used in geophysics and geodesy. | spherical harmonics, geophysics, geodesy, data analysis | Geodesy, Geophysics | Library | Documentation, Uses, and more |
| sibeliaz | mcc | 1.0 | SibeliaZ is a whole-genome alignment tool optimized for aligning many closely related genomes efficiently, identifying locally collinear blocks (LCBs) shared across the input sequences. It builds a compacted de Bruijn graph over all genomes to rapidly find conserved anchors, then constructs collinear blocks with its companion tool maf2synteny, scaling to dozens or hundreds of bacterial (or comparably sized) genomes far faster than traditional progressive aligners. The output is a multiple alignment in MAF format plus synteny-block coordinates that can feed comparative-genomics, rearrangement, and pan-genome analyses. Typical usage runs the aligner over a multi-FASTA of genomes and then performs the block-construction step. Version 1.2.5 is used in microbial and other closely-related-genome studies where whole-genome multiple alignment and synteny-block identification are needed at scale. | bioinformatics, genomics, structural variation, genomic analysis | Bioinformatics, Genomic Analysis | Analysis Tool | Documentation, Uses, and more |
| signalp | mcc, lcc | 1.0 | SignalP predicts the presence and precise location of signal-peptide cleavage sites in amino-acid sequences, distinguishing secretory signal peptides (which target proteins for the secretory pathway) from the mature protein and identifying where the signal peptidase cleaves. Version 4.1 uses neural networks trained on experimentally validated signal peptides and, given one or more protein sequences in FASTA, reports for each whether a signal peptide is predicted along with per-position scores (C-, S-, and Y-scores) and the most likely cleavage position; an organism-group option selects eukaryotes, Gram-positive, or Gram-negative bacteria, since signal-peptide characteristics differ between them, and output can be produced as a summary table or with graphical plots. It is a standard tool in secretome prediction and protein-annotation pipelines. SignalP is licensed academic software from DTU, and this deployment installs version 4.1 from a locally staged archive under the site's licensing arrangements. | bioinformatics, protein prediction, signal peptides | Bioinformatics, Molecular Biology | Prediction Tool | Documentation, Uses, and more |
| sigprofilerextractor | lcc | 1.0 | SigProfilerExtractor, from the Alexandrov Lab, performs de novo extraction of mutational signatures from catalogs of somatic mutations, a core method in cancer genomics for uncovering the mutational processes operating in tumors. It applies non-negative matrix factorization (NMF) with extensive bootstrapping and stability analysis to decompose a mutation-count matrix into a set of signatures and their per-sample activities, automatically evaluating a range of candidate signature counts and selecting a solution that balances reconstruction accuracy against reproducibility. It supports the standard SBS (single-base substitution, e.g. SBS-96), DBS (doublet), and ID (indel) mutational contexts and can take input directly from VCFs (via the companion SigProfilerMatrixGenerator), a mutation matrix, or MAF-like tables. It produces extracted de novo signatures, sample activities, stability/selection plots, and optional decomposition against the COSMIC reference signatures. | bioinformatics, cancer genomics, mutation analysis | Bioinformatics, Biochemistry and Molecular Biology | Command-line tool | Documentation, Uses, and more |
| simpleitk | lcc | 1.0 | SimpleITK is a simplified programming interface to the Insight Toolkit, providing image analysis and medical image processing capabilities through a concise API. It reads and writes the formats common in medical imaging, including NRRD, NIfTI, MetaImage and DICOM series, and preserves the spatial metadata that distinguishes medical images from ordinary arrays: voxel spacing, origin and direction cosines. Beyond input and output it supplies resampling and interpolation, registration, filtering, thresholding, morphology and connected-component labelling, and it converts to and from NumPy arrays so that images can move freely between it and the wider scientific Python stack. Note that the array returned to NumPy is indexed slice-first while the image itself is indexed column-first, a distinction that matters when combining the two. Here it underpins radiomic analysis, handling the reading of image volumes and label maps and the resampling needed to bring a mask into the geometry of its image. | Documentation, Uses, and more | |||
| singlem | mcc | 1.0 | SingleM, version 0.18.0, profiles the taxonomic composition and diversity of shotgun metagenomes (and metatranscriptomes) directly from reads by targeting conserved regions within a set of ~60 universal single-copy marker genes, without requiring assembly or reference genomes for the organisms present. Because it operates on the reads themselves, it can detect and quantify lineages absent from reference databases and estimate community richness even for the uncultured 'microbial dark matter'. Its companion Lyrebird mode targets viruses, and the associated 'microbial fraction' workflow (via its condense/read_fraction steps) estimates what proportion of a metagenome is prokaryotic. A basic profile is generated by running its pipe step on paired reads to produce a profile table, and profiles across many samples can be combined and summarized (alpha/beta diversity, appraisal against assembled genomes). It underpins large-scale surveys such as the Sandpiper resource of global metagenome taxonomic profiles. | Microbial Genomics, Metagenomics, Genome Assembly, Variant Calling | Genomics, Biological Sciences | Bioinformatics | Documentation, Uses, and more |
| singularity | lcc | N/A | Apptainer, formerly known as Singularity, is an open-source container platform designed to create portable and reproducible environments for scientific computing. It allows users to package applications, their dependencies, and data in a single container that can be run consistently across different computing environments. Description Source: https://apptainer.org/ |
Containerization, Scientific Computing, Hpc, Research | High-Performance Computing, Computer & Information Sciences, Engineering & Technology | Platform | Documentation, Uses, and more |
| singularity-devel | lcc | N/A | Singularity is a container platform designed for high-performance computing (HPC) environments, allowing users to create and run containers that encapsulate applications and their dependencies. | Containerization, HPC, Virtualization, Bioinformatics, Computational Chemistry | Computer Science, Software Engineering | Containerization Tool | Documentation, Uses, and more |
| sionlib | mcc, lcc | N/A | Scalable I/O Library for Parallel Access to Task-Local Files | I/O Interface, Scientific Data, Hdf5 Files | High-Performance Computing, Computer Science | Library | Documentation, Uses, and more |
| skera | mcc | 1.0 | skera (PacBio's pbskera, version 1.4.0 from bioconda) deconcatenates concatenated HiFi reads from Kinnex (formerly MAS-Seq) library preps, in which multiple cDNA or amplicon molecules are ligated together with known adapter sequences to maximize HiFi throughput. It scans each HiFi read for the Kinnex adapter sequences and splits the read at those positions into the constituent segmented reads (S-reads), effectively recovering the original individual molecules for downstream analysis. Its split operation takes the HiFi reads (BAM) and the appropriate array-adapter FASTA and produces a BAM of demultiplexed S-reads plus summary reports on segment counts and read-length distributions. It is the required first step of Kinnex full-length RNA (Iso-Seq) and 16S workflows before isoform or amplicon analysis. It preserves per-read metadata so segmented reads flow cleanly into Iso-Seq/lima downstream steps. | Documentation, Uses, and more | |||
| slepc | lcc | N/A | A library for solving large scale sparse eigenvalue problems | Eigenvalue Problems, Sparse Matrices, Parallel Computing | Numerical Linear Algebra, Other Mathematics | Numerical Computing | Documentation, Uses, and more |
| slow5tools | mcc | 1.0 | slow5tools is a toolkit for converting and manipulating Oxford Nanopore raw-signal data in the SLOW5/BLOW5 format, an efficient, well-documented alternative to the HDF5-based FAST5 format that suffers from slow, poorly parallelizable access on HPC and shared filesystems. BLOW5 is the compact binary form, and slow5tools provides subcommands to convert FAST5 to and from (B)SLOW5, merge and split files, index them for random access, and inspect or extract specific reads. Because SLOW5 supports efficient multi-threaded reading, it dramatically speeds up re-basecalling and signal-level analysis (e.g. with the slow5-enabled buttery-eel/f5c tools). A common step converts a directory of FAST5 files into BLOW5 and then merges the results. Version 1.3.0 is used to make large Nanopore signal datasets faster and cheaper to process on clusters. | bioinformatics, nanopore sequencing, data processing | Bioinformatics, Genomics | Command-line tools | Documentation, Uses, and more |
| slurm | mcc | N/A | Slurm is an open-source cluster resource management and job scheduling system that strives to be simple, scalable, portable, fault-tolerant, and interconnect agnostic. Slurm currently has been tested only under Linux. Description Source: https://github.com/SchedMD/slurm |
Workload Manager, Cluster Management, High-Performance Computing | Computer Science, Engineering & Technology | Workload Manager | Documentation, Uses, and more |
| slurm-monitoring | mcc | N/A | This module loads seff and reportseff to display SLURM usage stats. | Documentation, Uses, and more | |||
| snakemake | mcc | 1.0 | Snakemake is a Python-based workflow management system for building reproducible and scalable data-analysis pipelines, especially popular in bioinformatics. Workflows are defined in a Snakefile as a set of rules, each declaring input and output files plus the shell command, script, or Python code that produces the outputs; Snakemake infers the dependency DAG from filename patterns and wildcards and re-runs only the steps whose inputs changed. It transparently scales the same workflow from a laptop to clusters and clouds by mapping jobs onto schedulers (SLURM, SGE, etc.) or Kubernetes, supports per-rule software isolation via conda environments and containers, and provides features like threads/resources declarations, checkpoints, benchmarking, and report generation. Runs can execute locally across several cores or, for larger deployments, use conda environments or a cluster profile. This is version 6.12.3 from bioconda. | Workflow Management, Reproducibility, Traceability, Bioinformatics, Computational Biology, Data Science | Bioinformatics, Other Computer & Information Sciences | Software | Documentation, Uses, and more |
| snap | lcc | 1.0 | SNAP (Semi-HMM-based Nucleic Acid Parser) is an ab initio gene-prediction program that identifies protein-coding gene structures — exons, introns, splice sites, start and stop codons — directly from genomic DNA sequence using a hidden-Markov-model-based approach. It is trainable on a target organism: given a set of confirmed gene models it produces an HMM parameter file that captures the species' coding statistics and splice-site signals, which SNAP then applies to predict genes on new sequence. It is commonly used as one of the ab initio predictors within larger annotation pipelines such as MAKER, contributing gene models that are combined with evidence from RNA-seq and protein alignments. Input is a genome FASTA plus an HMM file, and output is gene predictions in a ZFF or GFF-style format. This is bioconda snap version 2013_11_29. | Network Analysis, Large Networks, Graph Algorithms | Computer Science, Computer & Information Sciences | Tool | Documentation, Uses, and more |
| snappy | mcc, lcc | N/A | Snappy is a compression/decompression library. It does not aim for maximum compression, or compatibility with any other compression library; instead, it aims for very high speeds and reasonable compression. Description Source: https://github.com/google/snappy |
Compression, Big Data, Data Processing | Computer & Information Sciences | Data Processing Tool | Documentation, Uses, and more |
| sniffles | mcc | 1.0 | Sniffles (version 2.2) is a fast, accurate structural variant (SV) caller designed for long-read sequencing data from PacBio and Oxford Nanopore instruments. Operating on coordinate-sorted BAM/CRAM alignments (typically produced by minimap2 or NGMLR), it detects deletions, insertions, duplications, inversions, and translocations, including at repetitive and complex loci that short reads miss. Sniffles2 introduced greatly improved speed and low memory use, accurate genotyping, a mosaic/somatic calling mode, and a population-scale workflow in which per-sample .snf files are produced and then jointly merged for cohort-level SV calling. Output is a standard VCF with detailed SV type, length, breakpoint, and support annotations. Repeat-aware and multi-sample analyses are enabled through optional tandem-repeat and .snf outputs. It is a mainstay of long-read variant-discovery and pangenome/structural-genomics studies. | genomics, bioinformatics, structural variants, long-read sequencing | Bioinformatics, Biochemistry and Molecular Biology | Command-line tool | Documentation, Uses, and more |
| sniffles-2.5.3 | lcc | 1.0 | Sniffles is a structural-variant (SV) caller purpose-built for long-read sequencing data from PacBio and Oxford Nanopore platforms. It analyzes coordinate-sorted, indexed BAM alignments (typically from minimap2, pbmm2, or NGMLR) and detects the full spectrum of SVs — insertions, deletions, duplications, inversions, translocations/breakends — using split-read and within-read (CIGAR) signals rather than short-read discordant pairs. Sniffles2 (the 2.x line) is substantially faster and lower-memory than v1, supports mosaic/low-frequency SV detection, and introduces a population/joint-calling workflow via per-sample .snf files that can be merged for multi-sample cohorts. A single-sample run produces a per-sample VCF and can optionally emit a .snf that is later merged across samples for population calling. Output is a standard VCF with genotypes; an optional tandem-repeat annotation improves calling in repetitive regions. | Genomics, Bioinformatics, Structural Variants | Bioinformatics, Genomics | Command-line Tool | Documentation, Uses, and more |
| snippy | lcc | 1.0 | Snippy is a fast pipeline for haploid variant calling and core-genome alignment, designed primarily for bacterial and other microbial genomes. Given sequencing reads (paired or single FASTQ, or even contigs) and a reference genome in FASTA or annotated GenBank/GFF format, it aligns reads with BWA-MEM, calls SNPs and small indels with FreeBayes under a haploid model, and annotates the resulting variants against the reference features, producing per-sample VCF, tab-delimited, and consensus-FASTA outputs. Its companion command snippy-core combines multiple per-sample results into a whole/core-genome SNP alignment suitable for downstream phylogenetics across an outbreak or population. This is bioconda snippy 4.6.0. | Haploid Variant Calling, Core Genome Alignments, Bacterial Isolates | Bioinformatics, Biological Sciences | Genomic Analysis Tool | Documentation, Uses, and more |
| snp-dists | lcc | 1.0 | snp-dists is a small, fast command-line utility that computes a pairwise SNP-distance matrix from a FASTA multiple-sequence alignment, counting the number of differing positions between every pair of aligned sequences. It is widely used in microbial genomics and outbreak/phylogenetic epidemiology to quantify how closely related isolates are once their genomes have been aligned (for example, from a core-genome or reference-based alignment). The input is a single aligned FASTA (all sequences the same length); the output is a tab-separated symmetric distance matrix that can feed clustering, transmission analysis, or visualization. Options let you emit molecular (per-site) distances, transpose or produce long/molten output, control the CSV/TSV separator, and set the number of threads. This is version 0.8.2 from bioconda. | Genomics, Bioinformatics, Sequence Analysis, Phylogenetics | Genomics, Biological Sciences | Analysis Tool | Documentation, Uses, and more |
| snp-sites | lcc | 1.0 | snp-sites is a small, very fast C tool for extracting variant (SNP) positions from a whole-genome multiple sequence alignment in multi-FASTA format, commonly used in bacterial phylogenomics to reduce large core-genome alignments to just their informative sites. It scans every column of the alignment and retains only those that vary across the samples, dramatically shrinking the data fed into downstream phylogenetic inference tools like RAxML or IQ-TREE. It can emit the variable sites in several formats — a reduced multi-FASTA alignment, a VCF (with positions relative to the input alignment), or PHYLIP — and can optionally include monomorphic (constant) sites in the output (useful for tools like BEAST) or restrict output to columns containing only unambiguous ACGT bases. This is bioconda snp-sites 2.5.1. | Bioinformatics, Computational Biology, Genomics | Bioinformatics, Biological Sciences | Bioinformatics Tool | Documentation, Uses, and more |
| soapdenovo | lcc | 1.0 | SOAPdenovo2 is a de novo whole-genome assembler designed for short-read (Illumina) data, and is a memory-efficient successor to the original SOAPdenovo geared toward large genomes. It follows a de Bruijn graph approach with distinct stages (k-mer graph construction, contig assembly, paired-end read mapping, scaffolding across insert-size libraries, and gap closing) that can be run together or step-by-step. It is driven by a configuration file describing each read library's insert size, rank, and file paths, and produces contig and scaffold FASTA outputs. The choice of k-mer length and the accompanying GapCloser step materially affect contiguity. This is version 2.40. | genome assembly, bioinformatics, high-throughput sequencing | Bioinformatics, Genomics | Open-source | Documentation, Uses, and more |
| solote | mcc | 1.0 | SoloTE (Solo Transposable Elements) is a tool for quantifying transposable-element (TE) expression at single-cell resolution from single-cell RNA-seq data, a key challenge because TEs are highly repetitive and multi-mapping. It processes an aligned scRNA-seq BAM (with cell-barcode and UMI tags, e.g. from Cell Ranger/STARsolo) together with a TE annotation, and separates locus-specific from subfamily-level TE signal to build a combined gene-plus-TE expression matrix per cell. That augmented matrix can then be loaded into standard single-cell frameworks such as Seurat or Scanpy to study TE contributions to cellular heterogeneity. It runs on the BAM plus a TE BED annotation to emit the SoloTE output matrix; this is version 1.09. | Documentation, Uses, and more | |||
| sortmerna | lcc | N/A | SortMeRNA is a local sequence alignment tool for filtering, mapping and clustering. | Bioinformatics, Hpc Tools, Computational Software | Biological Sciences | Bioinformatics Tool | Documentation, Uses, and more |
| sourmash | mcc | 1.0 | sourmash, version 4.9.4, is a command-line tool and Python library for rapid, scalable comparison of DNA/RNA and protein sequences using MinHash and FracMinHash sketches. Rather than aligning sequences, it reduces genomes, reads, or metagenomes to compact signatures that approximate containment and Jaccard similarity, enabling fast all-vs-all comparison, database search, and taxonomic classification even at petabase scale. Core operations include sketching to build signatures, compare and search for similarity and containment queries against reference databases (e.g. GTDB, GenBank), and gather for metagenome decomposition — the minimum-set-cover step that identifies which reference genomes are present in a sample. Combined with the sourmash tax module, gather results yield taxonomic profiles. It works with k-mer sizes for nucleotide or protein space and integrates with prepared LCA/SBT/zip databases for large-scale search. | Bioinformatics, Genomics, Nucleotide Signatures, Minhash Sketches | Bioinformatics, Biological Sciences | Bioinformatics Tool | Documentation, Uses, and more |
| space ranger | mcc | 1.0 | Space Ranger is 10x Genomics' official analysis pipeline for Visium spatial gene-expression assays, tying transcriptomic measurements to their physical location within a tissue section. It takes raw Illumina sequencing (via its FASTQ-generation step or provided FASTQs) together with the brightfield or fluorescence microscope image of the capture area and its fiducial/spot alignment, then performs read alignment (using a STAR-based aligner), spot barcode and UMI processing, and generation of a spatially-resolved feature-barcode matrix. Its main counting step outputs filtered/raw expression matrices, tissue-detection and spot-alignment results, per-spot QC metrics, and a web_summary.html plus a .cloupe file for interactive exploration in Loupe Browser. This 4.1.0 release supports both the original spot-based Visium and newer slide formats. It is the standard entry point before downstream spatial analysis in Seurat, Scanpy, or squidpy. | Documentation, Uses, and more | |||
| spaceranger | mcc, lcc | 1.0 | Cell ranger tools . | Bioinformatics, Hpc Tools | Molecular Biology, Biological Sciences | Bioinformatics Tool | Documentation, Uses, and more |
| spack | mcc, lcc, ecc | N/A | Build and installation framework | Package Management, Software Installation, Dependency Management, Hpc, Scientific Computing | High Performance Computing, Other Computer & Information Sciences | Utility | Documentation, Uses, and more |
| spades | mcc, lcc | 1.0 | SPAdes (St. Petersburg genome Assembler) is a de novo assembly toolkit built around a multi-sized de Bruijn graph, originally developed for single-cell bacterial sequencing but widely used for standard isolate bacterial and other small genomes. Its spades.py driver runs an integrated pipeline of read error correction (BayesHammer), iterative assembly over several k-mer sizes, graph simplification, mismatch/indel correction, and scaffolding using paired-end and mate-pair information, producing contigs.fasta, scaffolds.fasta, and an assembly graph. It offers specialized modes such as single-cell (MDA) mode, metagenome mode, isolate, plasmid, and RNA modes, and a careful mode for reduced misassemblies on small genomes, and accepts paired-end, mate-pair, and long-read (Nanopore/PacBio) libraries for hybrid assembly. Version 3.15.3 is provided as its own conda environment exposed as a Singularity app in a CentOS 7 baseline bioinformatics container. | Genome Assembly, Bioinformatics, Computational Biology | Bioinformatics, Biological Sciences | Bioinformatics Tool | Documentation, Uses, and more |
| speech-ipa-pipeline | lcc | 1.0 | speech-ipa-pipeline chains the speech tools in this container into the workflow a phonetician actually needs: it takes a recording, finds the speaker turns, transcribes each turn into the International Phonetic Alphabet, and writes the result already aligned to the audio for annotation software. Turns come either from speaker diarization run in the same job or from an RTTM file produced earlier, and any turn longer than the thirty second Whisper window is split into shorter pieces at the quietest point near an even division so no speech is lost. Each piece is transcribed by ZIPA, by WhIPA, or by both, and the output is written as a tab separated table of segments with start and end times, as a Praat TextGrid, and as an ELAN annotation file, each carrying a reference tier and an identical editing tier per speaker and model so a human can correct the machine transcription in place. An edited table can be fed back in to regenerate the annotation files without re-running any model. | Documentation, Uses, and more | |||
| splicev | lcc | N/A | SpliceV is a bioinformatics visualization tool for RNA-sequencing data that plots canonical splice junctions and backsplice (circular RNA) junctions together with read coverage across a gene, helping researchers inspect alternative splicing and circRNA structure. On LCC it is installed as the conda environment SpliceV (module ccs/conda/SpliceV, activating /opt/ohpc/pub/libs/conda/env/SpliceV); the environment ships the SpliceV command along with helper scripts RNABP.py and find_circ_convert. | Documentation, Uses, and more | |||
| splishsplash | lcc | 1.0 | SPlisHSPlasH is an open-source C++ research library from the Interactive Computer Graphics group for physically based simulation of fluids using Smoothed Particle Hydrodynamics (SPH). It implements a broad collection of state-of-the-art pressure/incompressibility solvers — including WCSPH, PCISPH, PBF, IISPH, DFSPH (Divergence-Free SPH), and Projective Fluids (PF) — along with models for viscosity, surface tension, vorticity/micropolar effects, drag, elastic solids, and rigid-fluid coupling. Simulations are configured through JSON scene files that define the fluid blocks, boundary geometry, solver choice, and physical parameters, and are run with the bundled SPHSimulator executable, optionally with an OpenGL GUI or in headless/batch mode for HPC. It exports particle data (e.g. VTK, partio, or its own formats) for downstream rendering and analysis, includes tools for surface reconstruction and partio conversion, and this build additionally compiles the pysplishsplash Python bindings so scenes can be driven and stepped from Python. | Documentation, Uses, and more | |||
| sqanti3 | mcc, lcc | 1.0 | SQANTI3 is the standard quality-control, classification, and filtering tool for long-read transcriptomes (PacBio Iso-Seq and Oxford Nanopore), used to evaluate transcript models defined from long reads against a reference annotation. It classifies each transcript's splice junctions relative to the reference into structural categories (full-splice-match, incomplete-splice-match, novel-in-catalog, novel-not-in-catalog, etc.), and computes many descriptive attributes and QC attributes (junction support, coverage, RT-switching, non-canonical junctions, cage/polyA signals) to flag potential artifacts. It then produces a rich HTML/PDF QC report and offers a machine-learning or rules-based filtering step to remove likely-false isoforms. Inputs are a transcript GTF/FASTA plus reference annotation and genome, with optional short-read (SJ) and CAGE/polyA evidence. This build is v4.2 (with cDNA_Cupcake), providing separate QC and filtering stages. | bioinformatics, transcriptomics, RNA-seq, genomics | Bioinformatics, Transcriptomics | Analysis Tool | Documentation, Uses, and more |
| sqlite | mcc, lcc, ecc | N/A | SQLite is a C-language library that implements a small, fast, self-contained, high-reliability, full-featured, SQL database engine. Description Source: https://www.sqlite.org/ |
Relational Database, Sql, Database Management | Computer Science, Computer & Information Sciences | Library | Documentation, Uses, and more |
| squashfs | ecc | N/A | SquashFS is a compressed read-only filesystem with built-in compression providing space-efficient storage. It is commonly used in embedded systems, live CDs, and Linux distributions to compress filesystem images and reduce storage space requirements while maintaining efficient access to files. | Filesystem, Compression, Read-Only | Computer Science, Computer & Information Sciences | Filesystem | Documentation, Uses, and more |
| squashfuse | ecc | N/A | Documentation, Uses, and more | ||||
| squigualiser | mcc | 1.0 | Squigualiser is a visualization tool for Oxford Nanopore sequencing data that renders the raw electrical signal ('squiggle') aligned against the corresponding base-called sequence, so users can inspect how individual bases and k-mers correspond to segments of the current trace. This signal-to-base visualization is valuable for methylation and modified-base analysis, basecaller/segmentation troubleshooting, and general quality inspection of nanopore data. It works with modern nanopore signal formats (SLOW5/BLOW5) together with a signal-to-read alignment (for example from tools like f5c/eventalign or move tables), and produces interactive HTML plots that can be zoomed and explored in a browser, including pileup-style views across multiple reads. The typical workflow prepares a signal alignment and then runs squigualiser's plot subcommands to generate the interactive figures. This is version 0.6.3 from bioconda. | Documentation, Uses, and more | |||
| sra-tools | mcc, lcc | 1.0 | The SRA Toolkit (version 2.11.0 from bioconda) is NCBI's official suite for downloading and converting data from the Sequence Read Archive, the primary public repository of raw sequencing reads. Its most-used utilities are prefetch, which downloads the compressed SRA (.sra) accession efficiently with resume support, and fasterq-dump/fastq-dump, which convert an accession or local .sra file into FASTQ, handling paired-end splitting, read filtering, and optional gzip. It also includes tools like vdb-config for configuring the local repository/cache and cloud access, and sam-dump for alignment retrieval. A typical workflow first prefetches an accession and then dumps it to FASTQ. It is the standard entry point for obtaining published sequencing datasets for reanalysis and is a prerequisite step in countless reproducible bioinformatics pipelines. | Bioinformatics, Hpc Tools, Sequence Analysis, Data Processing, High-Throughput Sequencing | Bioinformatics, Biological Sciences | Analysis Tool | Documentation, Uses, and more |
| sratoolkit | mcc, lcc | N/A | sratoolkit software. | bioinformatics, genomics, data processing | Biology, Genomics | Command-line tool | Documentation, Uses, and more |
| sshfs | mcc, ecc | N/A | SSHFS (SSH File System) is a filesystem client based on SSH (Secure Shell) that allows you to mount remote directories over a secure connection. | File System, SSH, Remote Access, Linux | Applied Computer Science, Computer Science | File System Client | Documentation, Uses, and more |
| stacks | mcc, lcc | 1.0 | Stacks is a software pipeline for analyzing restriction-site associated DNA sequencing (RAD-seq) data, designed for population genomics, phylogeography, and genetic-map construction in both model and non-model organisms. It assembles short reads into loci de novo (or against a reference), identifies SNPs within and across individuals, and builds a catalog of loci and alleles across a population, then exports genotypes and population statistics. The core programs include process_radtags for demultiplexing and quality-filtering raw reads, ustacks/cstacks/sstacks (or the reference-based gstacks) for building and matching stacks, and populations for computing summary statistics (heterozygosity, FST, nucleotide diversity) and exporting to formats like VCF, Structure, GenePop, and PLINK. The denovo_map and ref_map wrapper scripts run the full workflow end-to-end. This is version 2.65 from bioconda. | Software, Computational Biology, Bioinformatics | Genomics, Biological Sciences | Analysis Tool | Documentation, Uses, and more |
| star | mcc, lcc | 1.0 | STAR (Spliced Transcripts Alignment to a Reference) is a very fast, splice-aware aligner designed for RNA-seq, using a sequential maximal-mappable-prefix seed-search in an uncompressed suffix array followed by seed clustering and stitching to map reads across splice junctions and detect chimeric/fusion transcripts. It runs in two stages: genome index generation from a FASTA plus a GTF annotation, then alignment of the reads against that index. It can emit sorted BAM directly, per-gene read counts (via its gene-count quantification mode), and transcriptome-space alignments for RSEM/Salmon quantification, and supports a two-pass mode for improved novel-junction detection. Its speed comes at the cost of high memory (indexes for a mammalian genome need ~30GB RAM). Version 2.7.5c is a widely used front-end for RNA-seq expression and fusion-detection workflows. | Software, RNA-Seq, Read Aligner, Splice Junctions, Alternative Splicing | Genetics, Biological Sciences | Alignment Tool | Documentation, Uses, and more |
| starfish | mcc | 1.0 | starfish is a specialized bioinformatics toolkit for the de novo annotation and comparative analysis of Starships, a class of giant mobile genetic elements (large cargo-carrying transposons) found in fungal genomes. It provides a modular pipeline that first identifies the tyrosine-recombinase (captain) genes that mark Starship elements, then delimits element boundaries, groups elements into families and haplotypes, and analyzes their insertion sites and cargo-gene content across many genomes. It builds on standard genome-annotation inputs (assemblies, gene annotations, and protein/HMM databases) and orchestrates helper tools for gene finding, homology search, and orthology, producing element coordinates, family classifications, and visualizations of Starship distribution and synteny. It is organized into subcommands grouped into gene-finding, element-annotation, and region/visualization modules. Version 1.0.0 (installed via the egluckthaler conda channel on Python 3.8) is packaged as a per-tool conda environment in a Rocky 8 multi-tool bioinformatics container that also includes source builds of Dsuite, BayeScan, and blobtools. | Documentation, Uses, and more | |||
| stitch | mcc | 1.0 | STITCH (Sequencing To Imputation Through Constructing Haplotypes) is an R package by Robert Davies for genotype imputation and haplotype phasing directly from low-coverage whole-genome sequencing data, without requiring an external reference haplotype panel. It uses an expectation-maximization algorithm over a hidden Markov model of ancestral haplotypes to jointly estimate haplotypes and impute genotypes across many samples, making accurate genotyping feasible at very low sequencing depth (e.g. ~1x), which dramatically lowers the cost of population-scale studies. This is particularly valuable in non-human and non-model organisms where reference panels do not exist. It takes aligned BAM/CRAM files and a list of target positions and is run from R via its STITCH function with parameters such as the number of ancestral haplotypes K, the number of EM generations, the chromosome, and the position list, outputting a phased, imputed VCF. Version 1.6.10, built here as an R package on an R/devtools toolchain. | Documentation, Uses, and more | |||
| stringtie | mcc, lcc | 1.0 | StringTie is a fast and highly efficient transcript assembler that reconstructs full-length transcripts, including novel isoforms, from spliced RNA-seq alignments and quantifies their expression. Given a coordinate-sorted BAM from a splice-aware aligner such as HISAT2 or STAR, it builds a splice graph for each gene locus and applies a network-flow algorithm to assemble and simultaneously estimate the abundance of the isoforms that best explain the read coverage. It can perform reference-guided assembly with a GTF annotation, a merge step to build a unified set of transcripts across many samples, and produce Ballgown-ready tables or a per-transcript abundance matrix (via the prepDE script) for downstream differential-expression analysis. Version 2.1.7 is a core component of the HISAT2-StringTie-Ballgown RNA-seq workflow and is also used for long-read transcript assembly. | RNA-Seq, Transcript Assembly, Transcript Quantification | Biological Sciences | Transcriptomics Tool | Documentation, Uses, and more |
| structure | mcc, lcc | 1.0 | STRUCTURE is a classic population-genetics program that uses a Bayesian, Markov-chain-Monte-Carlo model-based clustering approach to infer population structure from multi-locus genotype data and to probabilistically assign individuals to a specified number K of ancestral populations. Under an admixture model it estimates each individual's fractional membership in each cluster, and it can be run across a range of K values to help infer the most likely number of populations, detect migrants and admixed individuals, and study hybrid zones. Input is a genotype matrix (a specifically formatted table of individuals, loci, and alleles) together with parameter files (mainparams and extraparams) controlling burn-in length, MCMC iterations, and model options; output consists of per-individual ancestry proportions and cluster allele frequencies, commonly post-processed with tools like CLUMPP and DISTRUCT. This is bioconda structure version 2.3.4. | Population Genetics, Population Structure, Genotype Analysis, Data Analysis | Ecology, Biological Sciences | Analysis Tool | Documentation, Uses, and more |
| structure_threader | mcc | 1.0 | Structure_threader (version 1.3.11 from PyPI) is a Python wrapper that parallelizes and orchestrates runs of population-structure inference programs, dramatically reducing wall-clock time by distributing the many independent runs across CPU cores. It drives STRUCTURE, fastStructure, MavericK, and ALStructure, automating the sweep over multiple K values (number of ancestral populations) and replicate runs, and it aggregates results and produces publication-ready plots including per-individual admixture bar plots and (for spatial data) interpolation maps. It also helps select the best-supported K, exposing separate run and plot stages for executing the analyses and then rendering the figures. It is popular in conservation and evolutionary genetics for making STRUCTURE-style analyses tractable on modern multicore hardware and for standardizing their visualization. | Documentation, Uses, and more | |||
| structures | lcc | N/A | ANSYS Structures (Mechanical / MAPDL) 25R2 is a finite element analysis suite for structural and mechanical engineering simulation. It solves static, dynamic, thermal, modal, buckling, and nonlinear problems on solid and shell models, supporting stress, deformation, fatigue, and contact analysis. It is used broadly in mechanical, civil, and aerospace engineering design and research. | Documentation, Uses, and more | |||
| studio | ecc | 1.0 | This 'Studio' is a combined Image and Video Studio delivered as an Open OnDemand interactive app, pairing ComfyUI with a JupyterLab front-end on a CUDA 12.1 PyTorch base. ComfyUI (from comfyanonymous/ComfyUI) is a powerful node/graph-based interface for Stable-Diffusion-style generative AI, letting users build image- and video-generation pipelines by wiring together nodes for model loading, prompting, sampling, latent operations, upscaling, ControlNet, and more — giving fine-grained, reproducible control over the diffusion workflow rather than a single prompt box. The bundled JupyterLab environment (with ipywidgets) provides a notebook interface whose helper code can drive the local ComfyUI instance programmatically, enabling scripted or batch generation alongside the visual graph editor. Users launch it from the cluster's OOD portal to get GPU-backed image and video generation. This build is from the ComfyUI GitHub repository plus pip-installed jupyterlab and ipywidgets. | Documentation, Uses, and more | |||
| subread | lcc | N/A | Subread carries out high-performance read alignment, quantification and mutation discovery. It is a general-purpose read aligner which can be used to map both genomic DNA-seq reads and DNA-seq reads. It uses a new mapping paradigm called seed-and-vote to achieve fast, accurate and scalable read mapping. Subread automatically determines if a read should be globally or locally aligned, therefore particularly powerful in mapping RNA-seq reads. It supports INDEL detection and can map reads with both fixed and variable lengths. | Ngs, Bioinformatics, Sequencing, Alignment | Bioinformatics, Biological Sciences | Ngs Analysis | Documentation, Uses, and more |
| superlu | lcc | N/A | A general purpose library for the direct solution of linear equations | Linear Algebra, Sparse Matrix, Direct Solver, Computational Science | Applied Mathematics, Mathematics | Solver Library | Documentation, Uses, and more |
| superlu-dist | mcc | N/A | SuperLU is a set of routines for solving large, sparse, nonsymmetric systems of linear equations. The 'dist' version is designed for distributed memory systems. | linear algebra, sparse matrices, parallel computing, HPC | Numerical Analysis, Applied Mathematics | Library | Documentation, Uses, and more |
| superlu_dist | lcc | N/A | A general purpose library for the direct solution of linear equations | Linear Solver, Sparse Systems, Parallel Computing | Applied Mathematics, Mathematics | Library | Documentation, Uses, and more |
| svim-asm | lcc | 1.0 | SVIM-asm is a structural-variant caller that detects SVs by comparing one or two genome assemblies against a reference, rather than from raw reads. It takes assembly-to-reference alignments (typically produced with minimap2, which is bundled in this environment along with samtools) and analyzes alignment breakpoints and gaps to call deletions, insertions, inversions, tandem and interspersed duplications, and translocations at base-pair resolution. It supports both haploid mode (single assembly) and diploid mode (two haplotype-resolved assemblies), the latter enabling genotyped, phase-aware SV calls. The input is a coordinate-sorted BAM of the assembly aligned to the reference plus the reference FASTA, and the output is a standard VCF of structural variants. A typical workflow aligns the assembly with an asm5-preset minimap2 run and then calls variants in haploid or diploid mode. This is version 1.0.2 from bioconda. | Documentation, Uses, and more | |||
| sweepfinder2 | mcc | 1.0 | SweepFinder2 is a population-genomics program for detecting and precisely localizing recent positive selection (selective sweeps) along a chromosome. It implements a composite-likelihood ratio (CLR) test that compares, at each test site, the observed spatial pattern of the site-frequency spectrum around that site against the genome-wide background spectrum, flagging the characteristic skew toward rare and high-frequency-derived alleles that a hitchhiking sweep produces. Version 2 improves on the original SweepFinder by allowing the background spectrum to be supplied directly, by incorporating a user-provided recombination map, and by optionally using invariant/substitution sites to increase power and reduce false positives from demography. Input is an allele-frequency file (and optionally a recombination file and a precomputed spectrum); it outputs CLR values and the estimated sweep parameter alpha at a grid of genomic positions, which are then scanned for peaks. Typical usage first computes the empirical background spectrum and then runs the sweep scan across the grid. Version 1.0 (the bioconda build number) is exposed as a conda app in a multi-domain genomics/phylogenetics container. | genomics, population genetics, bioinformatics, selective sweeps | Bioinformatics, Genomics | Analysis Tool | Documentation, Uses, and more |
| swig | mcc, lcc | N/A | SWIG (Simplified Wrapper and Interface Generator) is a software development tool for building scripting language interfaces to C and C++ programs. SWIG simplifies development by largely automating the task of scripting language integration--allowing developers and users to focus on more important problems. \r Description Source: https://www.swig.org/Doc4.2/SWIGDocumentation.html |
Software Development, Programming Tool, Scripting Languages, Interface Generator | Computer Science, Computer & Information Sciences, Software Engineering, Systems & Development | Interface Generator | Documentation, Uses, and more |
| syny | mcc | 1.0 | SYNY (version 1.2 from bioconda) is a pipeline for detecting and visualizing genome collinearity and synteny across multiple genomes, based on both protein and nucleotide sequence alignments. It identifies conserved gene order and homologous blocks between genomes, then generates a range of publication-quality visualizations, including linear/ribbon plots and circular Circos-style figures, to illustrate rearrangements, inversions, and conserved segments between species or strains. It takes annotated genomes (sequences plus annotations) as input, performs the alignments and clustering of collinear blocks internally, and outputs both the synteny block tables and the rendered figures. It is designed to make comparative structural-genomics analysis and its presentation straightforward from a single command. In this container it sits among a comparative/structural-genomics and population-structure toolset (NGenomeSyn, plotsr, SyRI, MCScanX, CLUMPP, KMC, vg), reflecting its role in cross-genome comparison workflows. | Documentation, Uses, and more | |||
| syri | mcc | 1.0 | SyRI (Synteny and Rearrangement Identifier) is a tool for comparative genomics that compares two chromosome-level genome assemblies to comprehensively identify their structural differences and local sequence variations. Starting from a whole-genome alignment between the two assemblies (typically produced with a mapper such as minimap2 or nucmer and provided as coords/BAM/PAF), it first determines the largest set of syntenic (collinear) regions and then classifies everything outside them as structural rearrangements — inversions, translocations, duplications, and their combinations — while within all aligned regions it annotates local variation such as SNPs and indels. Outputs are tabular files enumerating syntenic blocks, structural rearrangements, and sequence variants, which can be visualized with the companion plotsr tool. It requires both genomes to be at chromosome scale and typically consumes a nucmer-derived coordinate table together with the reference and query FASTAs. This is bioconda syri version 1.7.0. | Network Analysis, Visualization, Complex Networks | Computer & Information Sciences | Network Analysis Tool | Documentation, Uses, and more |
| sz | mcc, lcc | N/A | The sz modulefile defines the following variables:\r TACC_SZ_DIR, TACC_SZ_LIB, TACC_SZ_INC |
Documentation, Uses, and more | |||
| talon | mcc | 1.0 | TALON is a technology-agnostic pipeline for identifying and quantifying known and novel transcript isoforms from long-read RNA sequencing (PacBio Iso-Seq or Oxford Nanopore cDNA/direct-RNA), version 5.0. Working from long reads aligned to a reference genome, it annotates each read against a reference transcriptome and classifies its splice/isoform structure — known, ISM, NIC, NNC, antisense, intergenic, etc. — tracking novel isoforms across a whole dataset in a SQLite database so results are consistent across samples. Its typical workflow labels reads to flag internal priming, initializes a database from a GTF annotation, annotates SAM alignments and populates the database, then exports transcript-level count matrices and a custom GTF of observed (including novel) isoforms. It is part of the ENCODE long-read RNA-seq analysis toolkit and is used for isoform discovery and differential-isoform analyses. | Long-Read RNA Sequencing, Transcript Isoform Analysis, Gene Expression Quantification | Bioinformatics, Biological Sciences | Bioinformatics Tool | Documentation, Uses, and more |
| tama | mcc, lcc | 1.0 | TAMA (Transcriptome Annotation by Modular Algorithms) is a collection of Python tools for processing long-read (PacBio Iso-Seq / Oxford Nanopore) transcript alignments into transcript-model annotations, distributed from the GenomeRIK repository and run here with conda-provided biopython and pysam. Its two central modules are a collapse module, which collapses redundant mapped long reads (from a sorted BAM against a reference genome) into non-redundant transcript models while flagging potential artifacts such as internal priming and low-quality splice junctions, and a merge module, which merges transcript models from multiple samples or sources into a unified annotation with configurable tolerance for 5'/3' end and splice-junction variation. It is designed to give the user fine-grained, transparent control over how transcript ends and junctions are handled, in contrast to more black-box collapsing tools. Inputs are genome-aligned long reads (BAM) and a reference FASTA; outputs are BED12 transcript models plus supporting reports. It is commonly used to build isoform-level annotations and to characterize alternative splicing and transcript diversity. | Documentation, Uses, and more | |||
| tar | mcc, lcc, ecc | N/A | tar is a command-line utility used for archiving files and directories on Unix-based systems. It allows bundling multiple files and directories into a single archive file, often compressed using utilities like gzip or bzip2. tar is commonly used for backup, data transfer, and file compression tasks. | File Archiving, Compression, Terminal | Other, Computer & Information Sciences | File Management | Documentation, Uses, and more |
| tassel | lcc | 1.0 | canu software. | Bioinformatics, Genetic Analysis, Plant Biology | Genetics, Biological Sciences | Analysis Tool | Documentation, Uses, and more |
| tau | mcc, lcc | N/A | Tuning and Analysis Utilities Profiling Package | Software, Compiler, Hpc Tools | High-Performance Computing, Physical Sciences | Commercial | Documentation, Uses, and more |
| tbb | mcc, lcc, ecc | N/A | The Intel Threading Building Blocks (TBB) is a C++ library for parallel programming that enables developers to create multi-threaded applications efficiently. It offers high-level constructs for parallelism, task-based programming, and workload distribution, simplifying the development of scalable and performance-oriented applications. | Multicore Programming, Parallel Computing, C++ Library, High Performance Computing | Software Engineering, Computer & Information Sciences | Programming Library | Documentation, Uses, and more |
| tbb32 | mcc, lcc | N/A | Intel Threading Building Blocks (TBB) is a widely used C++ template library developed by Intel for parallel programming on multi-core processors. TBB provides higher-level abstractions for parallelism, making it easier to write code that can take advantage of multicore processors. | Parallel Programming, Multi-Core Processors, C++ Template Library, Intel | Computer & Information Sciences | Library | Documentation, Uses, and more |
| tcl | ecc | N/A | Tcl (Tool Command Language) is a powerful dynamic programming language suitable for a wide range of uses, including web and desktop applications, networking, administration, and testing. Description Source: https://www.tcl.tk/ |
Scripting Language, Dynamic Programming Language, Cross-Platform, Scripting | Software Engineering, Computer & Information Sciences | Language Interpreter | Documentation, Uses, and more |
| tcsh | mcc | N/A | tcsh is an enhanced version of the C shell (csh) with interactive command-line features for Unix-based systems. It provides a powerful shell with scripting capabilities, line editing, and command history functionalities. tcsh is known for its robust scripting language and interactive shell environment. | Command-Line Interpreter, Shell Scripting, Unix-Like Operating Systems | Operating Systems, Computer & Information Sciences | Language Interpreter | Documentation, Uses, and more |
| telseq | lcc | N/A | TelSeq is a C++ bioinformatics tool that estimates telomere length directly from whole-genome (or exome) shotgun sequencing BAM files by counting reads carrying telomeric repeat motifs (TTAGGG). Deployed on LCC as conda module ccs/conda/telseq-0.0.2. Used in genomics and aging/cancer research to quantify telomere length at scale without dedicated experimental assays; validated against experimental TRF measurements (Ding et al., NAR 2014). Open source (GPL), GitHub zd1/telseq. | bioinformatics, genomics, telomere analysis | Bioinformatics, Genomics | Command-line tool | Documentation, Uses, and more |
| tess3r | mcc | 1.0 | TESS3r is an R package (installed from the bcm-uga/TESS3_encho_sen GitHub repository) for spatially explicit population-genetics analysis, jointly estimating ancestry coefficients and ancestral allele frequencies from genotype data while accounting for the geographic coordinates of samples. It implements the TESS3 algorithm, which combines matrix factorization with a spatial regularization that encourages geographically close individuals to have similar ancestry, yielding estimates of individual admixture proportions across a chosen number of ancestral populations K. Beyond ancestry estimation, it provides genome scans for genotype-population differentiation to identify candidate loci under spatially varying selection, and it offers tools to interpolate and map ancestry coefficients as smooth geographic surfaces for visualization. Inputs are a genotype matrix plus a two-column matrix of sample longitude/latitude; the main tess3 function returns ancestry (Q) and frequency (G) estimates and a cross-validation criterion for choosing K, with helper functions to produce barplots and geographic ancestry maps. It is used in landscape-genomics studies of population structure. | Documentation, Uses, and more | |||
| tetools | mcc, lcc | 1.0 | Dfam TE Tools (this build wraps the dfam/tetools:1.4 Docker image) is a curated, ready-to-run bundle of the major transposable-element (TE) discovery and annotation software along with their many interdependencies pinned to compatible versions. It packages RepeatModeler for de novo identification and modeling of repeat families and RepeatMasker for annotating and soft/hard-masking those repeats in genome assemblies, together with supporting tools such as RMBlast, TRF, RECON, RepeatScout, and the Dfam/RepBase-compatible library machinery. This curated bundling solves the notoriously fragile dependency chain of the RepeatMasker/RepeatModeler ecosystem, making TE analysis reproducible. A typical workflow builds a species-specific repeat library with BuildDatabase followed by RepeatModeler, then runs RepeatMasker against the genome using that library to produce masked FASTA plus GFF/out annotation tables quantifying repeat content by family and class. It is a standard preprocessing step before gene annotation, since unmasked repeats otherwise confound gene predictors. | Documentation, Uses, and more | |||
| texinfo | mcc, lcc, ecc | N/A | Texinfo is the official documentation format of the GNU project. It is used to create both online information and printed output from a single source. | Documentation, Markup Language, GNU, Open Source | Software Engineering, Other Computer and Information Sciences | Documentation Tool | Documentation, Uses, and more |
| texlive | mcc | N/A | Comprehensive TeX document production system | Typesetting, Document Preparation, LaTeX, Open Source | Computer and Information Sciences, Computer Science | Typesetting System | Documentation, Uses, and more |
| time | mcc | N/A | The `time' command runs another program, then displays information about the resources used by that program, collected by the system while the program was running. | Time, Measurement, Events, Synchronization | Other | Measurement & Timekeeping | Documentation, Uses, and more |
| tk | mcc, lcc | 1.0 | Tk is a graphical user interface toolkit that takes developing desktop applications to a higher level than conventional approaches. | Gui, Python, Toolkit, Desktop Applications | Software Engineering, Computer & Information Sciences, Computer Science, Systems & Development | Library | Documentation, Uses, and more |
| tmhmm | lcc | 1.0 | TMHMM is a classic and highly cited predictor of transmembrane helices in proteins, using a hidden Markov model that captures the characteristic architecture of membrane proteins—hydrophobic membrane-spanning segments flanked by cytoplasmic and non-cytoplasmic loop regions. Given one or more amino-acid sequences in FASTA format, it predicts the number and location of transmembrane helices and the overall inside/outside topology of the chain, reporting per-residue probabilities and a summary of predicted TM regions, plus optional per-residue posterior plots. It is commonly used to distinguish integral membrane proteins from soluble ones and to annotate membrane topology in proteome-scale studies. This is version 2.0c, installed here from a locally staged licensed archive (TMHMM requires an academic license from DTU). | bioinformatics, protein structure, transmembrane prediction | Bioinformatics, Protein Structure Prediction | Prediction Tool | Documentation, Uses, and more |
| tophat | lcc | 1.0 | TopHat is a spliced read aligner for RNA-seq data that maps reads to a reference genome and, crucially, discovers exon-exon splice junctions without relying on a pre-existing gene annotation. Built on top of the Bowtie/Bowtie2 short-read aligner, it first maps reads that align contiguously, then takes the initially unmapped reads and splits them to identify reads that span introns, thereby detecting novel splice junctions de novo (it can also be guided by a supplied GTF annotation). Its output is a coordinate-sorted BAM of spliced alignments plus BED files of the identified junctions, insertions, and deletions, which then feed downstream tools such as Cufflinks for transcript assembly and quantification (the classic Tuxedo pipeline). Note that TopHat is now a legacy/low-maintenance tool superseded by HISAT2 and STAR, but it remains in use for reproducing older analyses. This build is version 2.1.1 from Bioconda. | RNA-Seq, Splice Junction Mapper, Genomics | Bioinformatics, Biological Sciences | Genomic Analysis Tool | Documentation, Uses, and more |
| totalview | mcc, lcc | N/A | TotalView is a scalable and interactive debugger for parallel and high-performance computing (HPC) applications. It provides advanced debugging features to help developers analyze and optimize the performance of parallel programs running on clusters or supercomputers. | Debugging, Performance Analysis, Hpc, Big Data, Machine Learning | Computer Science | Development & Optimization Tools | Documentation, Uses, and more |
| tpmcalculator | mcc, lcc | 1.0 | TPMCalculator computes gene- and transcript-level expression in TPM (Transcripts Per Million) directly from RNA-seq BAM alignments and a genome annotation, without requiring a separate quantification step. It parses a GTF to define exon/intron structures, then counts reads over those features and normalizes for both feature length and sequencing depth to yield TPM values, also reporting raw read counts and per-region (exon-level) results. Because it works from coordinate-sorted BAMs, it fits naturally after a splice-aware aligner such as STAR or HISAT2. Inputs are the annotation GTF and one or more BAM files (with options for paired-end handling, read/quality filtering, and library strandedness); outputs are tab-delimited tables of gene and transcript TPM plus optional per-exon detail. This is version 0.0.4 from bioconda. | Bioinformatics, Computational Biology, RNA-Seq, Transcriptomics | Biological Sciences | Bioinformatics | Documentation, Uses, and more |
| transdecoder | lcc | 1.0 | TransDecoder identifies likely protein-coding regions within transcript sequences, most commonly transcripts assembled de novo from RNA-seq (for example by Trinity) or reconstructed by genome-guided assemblers. It works in two stages: a LongOrfs step extracts all open reading frames above a minimum length, and a Predict step scores those ORFs using a Markov model of coding versus non-coding sequence (log-likelihood/Fickett-style criteria) and selects the most probable coding regions, optionally retaining ORFs that have supporting homology evidence from BLASTp against a protein database or Pfam domain hits from hmmscan. Inputs are a transcripts FASTA (and optional gene-to-transcript mapping and homology-search results); outputs are predicted peptide, CDS, GFF3, and BED files describing the coding regions. Version 5.5.0 is packaged as a conda tool in a broad CentOS 8 bioinformatics container. | Bioinformatics, Computational Biology, Transcriptomics, Protein Prediction | Genetics, Biological Sciences | Tool | Documentation, Uses, and more |
| transrate | lcc | 1.0 | Transrate (v1.0.3 prebuilt Linux binary) is a tool for reference-free and reference-based quality assessment of de novo transcriptome assemblies, addressing the difficulty of evaluating assemblies where no genome or gold-standard annotation exists. It maps the input reads back to the assembled contigs and, from that evidence, computes per-contig and overall assembly scores capturing metrics such as base-level accuracy, structural correctness, chimerism, and coverage support, ultimately reporting the Transrate assembly score and an optimized subset of well-supported contigs. When a reference proteome or related transcriptome is supplied, it additionally reports comparative metrics via reciprocal best-hit analysis (using BLAST/CRB-BLAST). Inputs are the assembly FASTA plus left/right read FASTQ files (and optionally a reference FASTA); outputs include a CSV of per-contig metrics, a good.assembly.fasta of high-quality contigs, and a summary of assembly scores. Note this legacy 1.0.3 binary bundles its own dependencies but is unmaintained. | Transcriptome Assembly, Quality Assessment, Bioinformatics | Genetics, Biological Sciences | Bioinformatics Tool | Documentation, Uses, and more |
| transrate-tools | lcc | 1.0 | transrate-tools (version 1.0.0 from bioconda) is the compiled C++ helper component of the Transrate transcriptome-assembly evaluation package, providing the fast BAM-processing routines that Transrate relies on. Specifically it parses read alignments to a de novo transcriptome assembly and computes per-contig support metrics (such as read coverage, proper-pair mapping, and base-level agreement) that feed Transrate's contig and assembly quality scores. It is not typically run directly by end users but is invoked internally when Transrate scores an assembly against its mapped reads; its key input is a coordinate-sorted BAM of reads aligned to the assembled transcripts. Together with Salmon/SNAP alignments and Transrate's scoring, it enables reference-free assessment of RNA-seq assemblies and identification of well- versus poorly-supported contigs. It is bundled in the container to satisfy Transrate's dependency chain for transcriptome QC. | Documentation, Uses, and more | |||
| travis | lcc | 1.0 | TRAVIS (Trajectory Analyzer and Visualizer) is a free, general-purpose program for analyzing and visualizing trajectories from molecular-dynamics and Monte Carlo simulations, developed primarily in the Kirchner group. It reads trajectory files (such as XYZ and other common MD formats) and computes a very wide range of structural and dynamical properties, including radial and spatial distribution functions, coordination numbers, hydrogen-bond analysis, dipole moments, mean-square displacement and diffusion coefficients, velocity autocorrelation and vibrational spectra (power spectra, IR, and Raman), and domain/aggregate analyses. It runs as an interactive text-menu-driven command-line program that guides the user through selecting atoms and observables, and it can also generate volumetric and rendered visualization output. This build is the 210521 source release. TRAVIS is widely used in computational chemistry to extract quantitative physics from bulk-phase and solution MD trajectories. | Documentation, Uses, and more | |||
| trf | mcc | 1.0 | Tandem Repeats Finder (TRF) is a long-established program that locates and characterizes tandem repeats in DNA sequences without requiring prior knowledge of the repeat pattern, period size, or number of copies. It uses a probabilistic model of tandem repeats based on percent identity and indel frequencies between adjacent copies to detect approximate repeats, reporting each repeat's location, period size, copy number, consensus pattern, percent matches/indels, and an alignment score. Input is a FASTA sequence file, and TRF is driven by a set of alignment-weight, matching-probability, and score thresholds, producing a data (.dat) and/or HTML report of all detected tandem repeats. It is a core component of repeat-annotation pipelines. Here TRF is provided via the Dfam TE Tools (TETools) v1.9.5 curated transposable-element annotation environment, packaged from the upstream dfam/tetools:1.95 Docker image. | Documentation, Uses, and more | |||
| trilinos | lcc | N/A | A collection of libraries of numerical algorithms | Scientific Computing, Numerical Analysis, Parallel Computing, Computational Physics | Biology, Engineering & Technology | Library | Documentation, Uses, and more |
| trim-galore | mcc, lcc | 1.0 | Trim Galore (version 0.6.10 from bioconda) is a wrapper that combines Cutadapt for adapter and quality trimming with FastQC for pre- and post-trimming quality reporting, streamlining the read-cleaning step of a sequencing pipeline. It automatically detects common adapter sequences (Illumina, Nextera, small-RNA) or accepts a user-specified adapter, trims low-quality 3' bases by Phred score, and removes reads that fall below a length threshold after trimming. It has dedicated support for paired-end data (validating read pairs and optionally trimming to equal lengths) and specialized modes for bisulfite/RRBS libraries, including handling of the filled-in cytosines at fragment ends. It produces trimmed FASTQs plus trimming and FastQC reports, and is a de facto standard preprocessing tool in RNA-seq, WGS, and methylation workflows. | Ngs Data Processing, Quality Control, Adapter Trimming | Bioinformatics, Biological Sciences | Data Processing Tool | Documentation, Uses, and more |
| trimmomatic | mcc, lcc | 1.0 | Trimmomatic (bioconda 0.39) is a widely used, flexible read-trimming and quality-control tool for Illumina next-generation sequencing data, supporting both single-end and paired-end reads. It performs adapter clipping (including palindrome-mode detection of adapter read-through in paired reads via its ILLUMINACLIP step), sliding-window quality trimming, leading/trailing low-quality base removal, and length filtering, applied as an ordered pipeline of processing steps. For paired-end input it maintains read pairing, emitting separate paired and unpaired output files so that reads whose mate was discarded are retained appropriately. Inputs are FASTQ files (optionally gzip-compressed) plus an adapter FASTA; outputs are trimmed FASTQ files and a trimming log summarizing surviving reads. It is a standard first step before alignment or assembly. | Ngs Data Processing, Quality Control, Bioinformatics | Biology, Biological Sciences | Tool | Documentation, Uses, and more |
| trinity | mcc, lcc | 1.0 | Trinity is a widely used method for the de novo reconstruction of full-length transcripts from RNA-seq reads, particularly valuable for organisms lacking a reference genome. It assembles transcriptomes in three sequential modules — Inchworm (assembles unique contig sequences from k-mers), Chrysalis (clusters related contigs into de Bruijn graph components representing genes and their isoforms), and Butterfly (traces reads and read-pairs through each graph to resolve alternatively spliced isoforms and paralogs) — producing a FASTA of assembled transcripts. It supports strand-specific libraries, paired- and single-end reads, in-silico read normalization for large datasets, and genome-guided assembly using a coordinate-sorted BAM. The accompanying utility scripts support downstream steps such as abundance estimation (RSEM/salmon), differential expression, and assembly-quality assessment. This build is version 2.13.2 from Bioconda. | Bioinformatics, Computational Biology, Transcriptomics, RNA-Seq, De Novo Assembly | Biological Sciences | Bioinformatics Tool | Documentation, Uses, and more |
| trinityrnaseq | lcc | 1.0 | Trinity is a widely used tool for de novo reconstruction of transcriptomes from RNA-Seq reads, designed for organisms without a reference genome. It proceeds through three sequential modules: Inchworm (assembling linear contigs from k-mers), Chrysalis (clustering contigs into components representing genes and building de Bruijn graphs), and Butterfly (resolving individual full-length transcripts and alternatively spliced isoforms from those graphs). This official 2.12.0 Docker image also bundles supporting tools such as samtools and RSEM so users can align reads back to the assembly and quantify transcript/gene abundance in the same environment. It takes paired or single FASTQ reads and produces a Trinity.fasta of assembled transcripts plus a gene-to-transcript mapping. Downstream utilities support abundance estimation, differential expression, and assembly-quality assessment, making it an end-to-end de novo transcriptomics workhorse. | RNA-Seq, Transcriptome, Assembly, Bioinformatics | Bioinformatics, Biological Sciences | Assembly Tool | Documentation, Uses, and more |
| trinotate | lcc | N/A | Trinotate is a comprehensive annotation suite designed for automatic functional annotation of transcriptomes, particularly de novo assembled transcriptomes, from model or non-model organisms. | Bioinformatics, Computational Biology, Transcriptomics, Functional Annotation | Bioinformatics, Biological Sciences | Functional Annotation Tool | Documentation, Uses, and more |
| trnascan-se | lcc | 1.0 | tRNAscan-SE is the standard tool for detecting and annotating transfer-RNA (tRNA) genes in genomic DNA. It combines two fast pre-filter scanners (tRNAscan and an Aragorn/EufindtRNA-style search) with a highly specific covariance-model (Infernal) search that models the conserved secondary structure of tRNAs, achieving both high sensitivity and very low false-positive rates; version 2.x uses domain-specific covariance models for bacteria, archaea, and eukaryotes and reports the isotype (amino acid), anticodon, predicted secondary structure, and a score for each candidate, while flagging pseudogenes and introns. Input is a genome/sequence FASTA and output is a tabular list of predicted tRNA genes with coordinates and structures, along with optional secondary-structure and FASTA outputs. Users select the search domain (eukaryotic, bacterial, or archaeal) to match their input. Version 2.0.9 is provided as a conda tool in a Rocky 8 multi-environment container that also bundles quantum-computing and utility tools. | Bioinformatics, Genomics, Computational Biology, Sequence Analysis | Genomics, Biological Sciences | Bioinformatics Tool | Documentation, Uses, and more |
| trtools | mcc | 1.0 | TRTools is a Python toolkit developed by the Gymrek lab for working with tandem-repeat (short tandem repeat / STR and variable-number tandem repeat / VNTR) genotypes produced by specialized TR genotyping callers. It provides a set of command-line utilities that operate on VCF files emitted by tools such as HipSTR, GangSTR, ExpansionHunter, adVNTR, and PopSTR, harmonizing their differing conventions. Key commands include dumpSTR (per-call and per-locus quality filtering), mergeSTR (merging VCFs across samples), compareSTR (comparing call sets from different callers or the same caller across runs), qcSTR (generating quality-control plots), and statSTR (computing summary statistics like heterozygosity and allele frequencies). Inputs are bgzipped, indexed VCFs and outputs are filtered/merged VCFs plus QC reports and plots. This is bioconda trtools 6.0.1. | Python Library, Data Transformation, Tabular Data, Data Manipulation | Computer & Information Sciences | Library | Documentation, Uses, and more |
| ttt | mcc | 1.0 | TTT (Trivial Tangle Traverser) from the marbl group is a specialized utility for resolving repetitive "tangles" in genome assembly graphs produced by Verkko or hifiasm. Highly repetitive regions such as satellites and segmental duplications often collapse into complex knots of nodes and edges (tangles) in the assembly graph that automated assemblers cannot untangle, blocking telomere-to-telomere completion. TTT works on the GFA assembly graph together with GAF-format alignments of the long reads to that graph, using the read paths to find supported traversals through a tangle and thereby propose how the repeat should be resolved into a linear path. It is aimed at expert assembly-curation workflows for finishing near-complete genomes, complementing tools like verkko-fillet. Inputs are the GFA graph and GAF read-to-graph alignments; output is proposed traversal/resolution information the curator applies to the graph. It is distributed from the marbl GitHub repository as part of the T2T assembly toolchain. | Documentation, Uses, and more | |||
| ucsc-bedgraphtobigwig | mcc | 1.0 | bedGraphToBigWig is a UCSC Genome Browser utility (packaged as version 377) that converts a bedGraph coverage/signal file into the compact, indexed binary bigWig format used for efficient random-access display and analysis of genome-wide quantitative tracks. bigWig files can be served remotely and queried by region without loading the whole file, which is why they are the standard for coverage, ChIP-seq/ATAC-seq signal, and other continuous genomic data in genome browsers and tools like deepTools. It takes a position-sorted bedGraph together with a chromosome-sizes file (chromosome lengths, obtainable via fetchChromSizes for a given assembly) and writes an indexed bigWig. Overlapping bedGraph intervals are not allowed, so the input is typically produced by tools such as bedtools genomecov or deepTools. It is a small, fast conversion step within epigenomics and coverage-visualization pipelines. | genomics, bioinformatics, data visualization | Bioinformatics, Biological Sciences | Command-line tool | Documentation, Uses, and more |
| ucsc-fatotwobit | lcc | N/A | faToTwoBit is a UCSC Genome Browser command-line utility that converts FASTA sequence files into the compact, indexed .2bit binary format used by BLAT, the UCSC browser, and many downstream genomics tools. On LCC it is deployed as a Miniconda3 conda environment (ucsc-fatotwobit-377) that provides the faToTwoBit executable. It is a small but genuinely user-facing bioinformatics data-format converter used when preparing reference genomes for browser tracks and alignment pipelines. | Documentation, Uses, and more | |||
| ucsc_utils | mcc | 1.0 | The UCSC Genome Browser command-line utilities are a collection of small, fast standalone programs (the 'kent' tools) for converting and manipulating genomic sequence and annotation file formats, used ubiquitously as glue in genomics pipelines. Representative tools include faToTwoBit and twoBitToFa (FASTA <-> compact .2bit), faSize (sequence statistics), bedToBigBed and bedGraphToBigWig (creating indexed binary track files), bigWigToWig, liftOver (coordinate conversion between assemblies), and various sorting/filtering helpers. Each is a single-purpose binary run on the command line. In this environment they are supplied as part of the Dfam TE Tools (TETools) v1.9.5 environment packaged from the upstream dfam/tetools:1.95 Docker image, where they support the repeat-annotation workflow. | Documentation, Uses, and more | |||
| ucx | mcc, lcc, ecc | N/A | UCX is a communication library implementing high-performance messaging | Communication Library, High-Performance Computing, Distributed Computing | Computer Science | Communication Library | Documentation, Uses, and more |
| ukyaro427 | lcc | 1.0 | Documentation, Uses, and more | ||||
| ultra | mcc | 1.0 | uLTRA (pip package ultra-bioinformatics) is a splice-aware long-read alignment tool specialized for aligning cDNA/mRNA long reads (PacBio Iso-Seq and Oxford Nanopore cDNA/direct-RNA) to a genome with high accuracy at exon boundaries and small exons. It uses a maximal-exact-match seeding strategy (leveraging slaMEM/StrobeMap in this build) to find candidate loci and a database of known transcript exon/junction structures derived from an annotation, and it combines this with minimap2 to produce accurate spliced alignments, notably improving the correct placement of reads over small exons and precise splice-site alignment compared to alignment without annotation guidance. Inputs are a reference genome FASTA, a gene annotation (GTF) that is first indexed with the uLTRA index/prep step, and the long-read FASTQ/FASTA; the output is a spliced-alignment SAM/BAM. Its pipeline mode builds the index and aligns in one step, producing alignments suitable for isoform detection and quantification. It is used in long-read transcriptomics to get precise transcript alignments before isoform collapsing. | Documentation, Uses, and more | |||
| umf | ecc | N/A | Documentation, Uses, and more | ||||
| uncalled4 | mcc | 1.0 | UNCALLED4 is a toolkit for aligning and analyzing the raw electrical signal (squiggle) produced by Oxford Nanopore sequencers, rather than only the basecalled sequence. It performs signal-to-reference alignment — mapping the raw current measurements of each read onto a reference sequence — and stores these alignments in a compact, indexed format that supports a range of downstream signal-level analyses. A major application is the detection and comparison of DNA/RNA modifications (such as methylation) by examining deviations between observed and expected signal at reference positions across groups of reads, and it provides visualization and statistical comparison tools (e.g. per-reference-position current distributions and dotplots) for interrogating those signals. It builds on and extends the earlier UNCALLED real-time selective-sequencing work. Conceptually, an align stage produces signal alignments (using a reference, reads, and a pore model) and compare/refstats stages drive downstream analysis. This deployment installs uncalled4 version 4.1.0 from PyPI. | Documentation, Uses, and more | |||
| unzip | mcc, ecc | N/A | UnZip is a utility tool used for extracting and viewing files compressed in the ZIP format. It allows users to decompress ZIP archives, extract individual files, and preserve the directory structure of the compressed content. | File Extraction, Compression, Utility | Computer Science, Computer & Information Sciences | File Extraction | Documentation, Uses, and more |
| util-linux-uuid | mcc, lcc, ecc | N/A | util-linux-uuid is a command-line utility that allows users to generate universally unique identifiers (UUIDs) in various formats. UUIDs are 128-bit numbers used as identifiers for entities in computer systems with a high probability of being unique. | Uuid Generator, Command-Line Utility, Unique Identifier, Data Management | Computer Science, Computer & Information Sciences | Command-Line Tool | Documentation, Uses, and more |
| util-macros | mcc, lcc, ecc | N/A | util-macros is a collection of utility macros for C programmers to aid in simplifying common tasks and improving code readability. These macros are designed to enhance the efficiency and maintainability of C code by providing useful functions and utilities. | Programming, Development, C Programming, Utility Macros | Computer Science, Computer & Information Sciences | Utility | Documentation, Uses, and more |
| valgrind | mcc, lcc | N/A | Memory debugging utilities | Debugging, Profiling, Development Tool | Computer Science, Computer & Information Sciences | Programming Tool | Documentation, Uses, and more |
| vamb | mcc | 1.0 | VAMB (Variational Autoencoder for Metagenomic Binning) is a deep-learning metagenomic binning tool that groups assembled contigs into bins approximating individual genomes (MAGs). It encodes each contig's tetranucleotide (k-mer) composition together with its abundance profile across multiple samples into a shared latent space using a variational autoencoder, then clusters the latent representations, which lets co-abundance information across samples improve bin separation. Inputs are a contigs FASTA and per-sample abundance information (BAM alignments or precomputed depth), and outputs are cluster/bin assignments and binned FASTA files. It runs from the command line and benefits from GPU acceleration for the neural-network training. This is version 3.0.2. | bioinformatics, metagenomics, machine learning, deep learning, data analysis | Computational Biology, Bioinformatics | Open-source | Documentation, Uses, and more |
| variantbam | lcc | 1.0 | VariantBam, version 1.4.4a, is a tool from the Broad Institute for fast, rule-based filtering, subsampling, and extraction of reads from BAM/CRAM files, letting users massively reduce sequencing data to just the reads relevant to an analysis. Reads are kept or discarded according to a flexible rules engine that can combine genomic-region membership (from BED/VCF/interval lists) with read properties such as mapping quality, alignment flags, insert size, presence of indels/clips, mate mapping, and duplicate status. It also supports coverage-based subsampling to cap depth at a target level, which is useful for trimming ultra-high-coverage regions before variant calling or visualization. Rules can be supplied as a JSON file or a compact command-line rule string, producing a smaller BAM plus optional read-count statistics. It is valuable for building targeted datasets, reducing storage, and speeding downstream analyses. | Documentation, Uses, and more | |||
| vasp | mcc, lcc | N/A | vasp software. | Materials Modelling, Electronic Structure Calculations, Quantum-Mechanical Simulations, Dft Calculations, Molecular Dynamics | Chemistry, Physical Sciences, Engineering & Technology | Package | Documentation, Uses, and more |
| vaspkit | mcc, lcc | N/A | VASPKIT post-processing tool for VASP. | VASP, post-processing, visualization, density of states, band structure | Materials Science, Condensed Matter Physics | Post-processing tool | Documentation, Uses, and more |
| vcf2arlequindiploid | mcc | 1.0 | VCF2ArlequinDiploid is a lightweight converter script that bridges variant-calling output and population-genetics analysis by transforming diploid SNP genotypes stored in VCF format into the Arlequin project (.arp) input format. Arlequin expects a specific structured layout of populations and diploid haplotype/genotype data, and this tool automates that reformatting so users can avoid error-prone manual editing. It reads a standard VCF (per-sample GT fields), assigns individuals to populations, and writes the corresponding Arlequin data blocks. The typical workflow is to filter/normalize a VCF (e.g., biallelic SNPs), run the script to emit the .arp file, then load that file into Arlequin to compute statistics such as AMOVA, F-statistics, heterozygosity, and Hardy-Weinberg tests. It is a niche utility valuable to conservation/population geneticists who use the Arlequin ecosystem. | Documentation, Uses, and more | |||
| vcflib | mcc, lcc | 1.0 | vcflib is a C++ library and accompanying suite of command-line utilities for parsing, filtering, transforming, and annotating VCF (Variant Call Format) files produced by variant callers. It provides a lightweight VCF/genotype API for developers plus dozens of focused programs such as vcffilter (filter records by INFO/genotype expressions), vcfbreakmulti and vcfallelicprimitives (decompose multi-allelic and complex variants into canonical primitives), vcfleftalign (left-align indels), vcfstats, vcfintersect, and vcfrandomsample. It is frequently used as glue in variant-processing pipelines to simplify, subset, and harmonize VCFs before downstream analysis, and its decomposition and left-alignment tools are valuable for reconciling calls across callers. Records are read from files or streamed via stdin/stdout, making the tools easy to chain in shell pipelines. This is version 1.0.9 from bioconda. | bioinformatics, genomics, VCF, C++ | Bioinformatics, Genomics | Library | Documentation, Uses, and more |
| vcftools | mcc, lcc | 1.0 | VCFtools is a widely used toolkit for working with variant call format (VCF) files, providing a rich set of options for filtering, summarizing, comparing, and converting genetic variant data. Its main vcftools command applies filters (by region, position, quality, allele frequency, missingness, indels vs SNPs, sample subsets) and computes population-genetic statistics such as allele frequencies, per-site and per-individual missingness, heterozygosity, nucleotide diversity (pi), Tajima's D, Fst, Hardy-Weinberg equilibrium, and linkage disequilibrium. It can also convert VCF to other formats (e.g. PLINK, IMPUTE) and, via the companion vcf-tools/vcf-* Perl scripts, merge, concatenate, and compare VCFs. It reads plain or bgzipped VCF and outputs filtered VCFs and statistic tables. This is bioconda vcftools version 0.1.16. | Genetic Variation Analysis, Variant Call Format (Vcf), Population Genetics, Bioinformatics | Bioinformatics, Biological Sciences | Data Analysis Tool | Documentation, Uses, and more |
| vclust | mcc | 1.0 | Vclust (bioconda 1.3.1) is a fast, scalable tool for clustering viral (and other) genome sequences based on average nucleotide identity (ANI), designed to define species- and genus-level operational taxonomic units from very large sequence collections such as those produced by viral metagenomics. It computes pairwise ANI efficiently using a k-mer-based prefiltering step (Kmer-db) to avoid all-versus-all comparison of dissimilar sequences, followed by a Lempel-Ziv-based pairwise aligner (LZ-ANI) on candidate pairs, then applies configurable identity and coverage thresholds and a choice of clustering algorithms (via Clusty) to group genomes — for example using the ~95% ANI species and ~70% genus thresholds recommended for prokaryotic viruses by ICTV. The workflow proceeds conceptually in three stages: a prefilter stage to identify candidate pairs, an align stage to compute ANI, and a cluster stage to produce cluster assignments at the chosen thresholds. Inputs are a multi-FASTA of genomes; outputs are pairwise ANI tables and cluster/membership files. It is well suited to dereplicating and taxonomically organizing large sets of viral contigs from environmental samples. | Documentation, Uses, and more | |||
| vcontact2 | mcc | 1.0 | vConTACT2 (bioconda 0.11.3) is a tool for genome-based, reference-independent taxonomic classification of viral sequences, especially bacteriophages, using gene-sharing networks. It compares the predicted protein content of query viral genomes against each other and against a reference database (e.g., NCBI/ICTV prokaryotic viruses), computes protein clusters, and builds a network in which viruses are nodes connected by edges weighted by the number of shared protein clusters. It then applies clustering (ClusterONE) to that network to group genomes into viral clusters that approximate genus-level taxonomy, placing novel viruses relative to known references. Inputs are predicted proteins (from Prodigal) plus a gene-to-genome mapping file; outputs include the genome-by-genome clustering assignments, the network files, and summary tables, and the resulting network is commonly visualized in Cytoscape. It is a standard step in viral-metagenomics (viromics) classification workflows. | Documentation, Uses, and more | |||
| vcontact3 | mcc | 1.0 | vConTACT3 is a tool for the taxonomic classification of viral genomes (particularly bacteriophages and other prokaryotic viruses) through whole-genome gene-sharing network analysis. Rather than relying on single marker genes, it computes the shared protein content among viral genomes, builds a network in which genomes are nodes connected by the number of shared protein clusters, and clusters this network to define viral genome groups that approximate taxonomic ranks, integrating reference databases to propagate known taxonomy to unknown query genomes. Input is viral genome sequences or predicted proteins together with gene-to-genome mappings, and output includes network files, genome clustering assignments, and taxonomic predictions that can be visualized (e.g. in Cytoscape). It is widely used in viral metagenomics to place newly recovered viral contigs into a taxonomic framework. This is bioconda vcontact3 version 3.1.6. | Documentation, Uses, and more | |||
| velocyto | mcc, lcc | 1.0 | velocyto.py (version 0.17.17 from bioconda) implements RNA velocity analysis for single-cell RNA sequencing, estimating the future transcriptional state of individual cells by contrasting unspliced (nascent, intron-containing) with spliced (mature) mRNA counts. Its command-line component annotates aligned reads against gene models to separate reads into spliced, unspliced, and ambiguous categories, producing a loom file of the corresponding count matrices for each cell. Its run10x mode processes 10x Genomics Cell Ranger output directly (optionally with a repeat-mask GTF), while its general run mode handles arbitrary BAMs. The resulting loom file is then analyzed in Python (velocyto/scVelo) to compute velocity vectors and project them onto embeddings such as UMAP/t-SNE, revealing directional cell-state transitions and differentiation trajectories. It is a foundational tool in developmental and dynamics-focused single-cell studies. | bioinformatics,single-cell RNA-seq,RNA velocity,data analysis | Molecular Biology, Biological Sciences | Analysis Tool | Documentation, Uses, and more |
| velvet | mcc, lcc | 1.0 | Velvet is one of the original de Bruijn graph short-read de novo assemblers, designed for assembling short next-generation sequencing reads into contigs and, with paired-end information, scaffolds. It operates in two stages: a hashing step (velveth) hashes the input reads into k-mers and builds the initial data structures for a chosen k-mer length, and a graph step (velvetg) then constructs and simplifies the de Bruijn graph, resolves repeats, uses paired-end/insert-size information for scaffolding, and emits contigs. Key parameters include the hash (k-mer) length, expected coverage, and coverage cutoff for removing low-coverage (likely erroneous) nodes. It accepts FASTA/FASTQ input in various read categories (short, shortPaired, long) and outputs a contigs file plus assembly statistics. Version 1.2.10 is provided as its own conda environment/Singularity app in a CentOS 7 baseline bioinformatics container; it is best suited to smaller genomes given its memory profile. | Genomics, Bioinformatics, Sequencing, Assembly | Bioinformatics, Biological Sciences | Bioinformatics Tool | Documentation, Uses, and more |
| verifybamid2 | mcc | 1.0 | VerifyBamID2 estimates the level of cross-individual DNA contamination in a sequencing sample and helps detect sample swaps directly from aligned reads. It compares observed sequence read data at known polymorphic sites against a reference panel of population allele frequencies, simultaneously estimating the contamination fraction and the sample's genetic-ancestry coordinates so that ancestry differences do not bias the contamination estimate — an improvement over the original VerifyBamID. It works on BAM/CRAM files without requiring the sample's own array genotypes. Inputs are the alignment file, the reference FASTA, and a bundled resource panel (SVD-decomposed allele frequencies such as the 1000 Genomes-derived panels); the primary output is the estimated contamination value (FREEMIX). It is routinely used as a QC step in large sequencing pipelines. This container wraps the upstream griffan/verifybamid2 Docker image, adding only cluster mount points. | Documentation, Uses, and more | |||
| verkko | mcc | 1.0 | Verkko (version 2.0) is a hybrid genome-assembly pipeline from the T2T/marbl group designed to produce telomere-to-telomere, near-complete, and often haplotype-resolved assemblies. It combines the accuracy of PacBio HiFi reads with the length of Oxford Nanopore ultra-long reads: HiFi reads build an initial high-resolution assembly graph (via MBG), ultra-long ONT reads are aligned (GraphAligner) to resolve repeats and simplify the graph, and optional Hi-C or trio (parental k-mer) data phase the result into haplotypes. Verkko is implemented as a Snakemake pipeline that orchestrates all steps and can run on a single machine or distribute across a Slurm cluster. Inputs are HiFi and ONT read files (plus optional Hi-C/parental data) and outputs are assembly graphs (GFA) and phased contig/scaffold FASTA files. It is a leading tool for reference-quality and pangenome assembly projects. | Documentation, Uses, and more | |||
| verkko-2.2.1 | mcc | 1.0 | Verkko (bioconda 2.2.1) is a hybrid, largely automated pipeline for producing telomere-to-telomere (T2T), near-complete genome assemblies by combining the complementary strengths of accurate PacBio HiFi reads and ultra-long Oxford Nanopore reads. It first builds a multiplex de Bruijn graph from HiFi data to obtain accurate contigs, then uses the ultra-long ONT reads to resolve repeats and traverse tangles in the graph, and can integrate Hi-C or trio (parental k-mer) information for phasing into complete haplotypes. Internally it is orchestrated as a Snakemake workflow that chains graph construction (MBG), graph simplification, ONT-based path finding (GraphAligner), and consensus generation, so a single command drives the whole process. Inputs are HiFi and ultra-long ONT read sets (optionally plus Hi-C or parental reads); outputs are highly contiguous, often T2T-complete assembly FASTA files with an assembly graph. It is well suited to reference-grade assembly of individual genomes on HPC resources. | Documentation, Uses, and more | |||
| verkko-fillet | mcc | 1.0 | verkko-fillet (PyPI version 0.1.24, from the marbl group) is a post-Verkko curation framework that helps users inspect and manually correct telomere-to-telomere (T2T) genome assembly graphs. After Verkko produces an assembly, difficult regions often need expert curation — resolving remaining tangles, fixing misjoins, patching gaps, and verifying node paths — and verkko-fillet provides a Python/Snakemake-based interface that wraps Verkko together with a suite of assembly-QC and visualization tools to make this iterative graph editing tractable. This build installs the full upstream conda environment, including Verkko 2.2.1, samtools, minimap2, mashmap, gnuplot, and R/Bioconductor packages, so the container supports the complete inspect-edit-reassemble loop. It operates on Verkko's assembly graph (GFA) and supporting read-alignment data, letting a curator examine node/edge structure and apply corrections toward a finished genome. It is aimed at expert assembly-finishing workflows for reference-grade T2T projects, complementing tools like Verkko and TTT. | Documentation, Uses, and more | |||
| vg | mcc | 1.0 | vg (variation graphs), version 1.63.1, is a toolkit for building, manipulating, indexing, and mapping reads to genome variation graphs — pangenome references that encode multiple genomes and their variants as a graph rather than a single linear sequence. By representing known variation in the reference itself, vg reduces the reference bias that afflicts linear-genome alignment and improves mapping and genotyping accuracy in polymorphic or structurally variable regions. It provides subcommands to construct graphs from a reference plus VCF or from assemblies, build the indexes required for mapping (xg, GBWT, GCSA, or the newer giraffe indexes), align reads with the standard map or the faster giraffe mapper, and call variants; the graph and alignment formats are GFA and GAM. A typical flow proceeds from graph construction to indexing to read mapping to variant calling. It is a central tool in pangenomics for graph-based read mapping, variant genotyping, and structural-variant analysis across populations. | genomics, bioinformatics, variant analysis, graph theory | Bioinformatics, Genomics | Command Line Tool | Documentation, Uses, and more |
| vibrant | mcc | 1.0 | VIBRANT (Virus Identification By iteRative ANnoTation) is a tool for automated recovery, annotation, and functional characterization of bacterial and archaeal viruses (bacteriophages), including integrated prophages, from genomic and metagenomic sequence data. It uses neural-network classification of protein annotation signatures — combining HMM searches against Pfam, KEGG, and its curated VOG viral databases — to distinguish viral from microbial sequences and to assess genome quality/completeness, and it specifically identifies and quantifies auxiliary metabolic genes (AMGs) that reflect viral influence on host metabolism. Input is a nucleotide FASTA of scaffolds/contigs (or proteins), and outputs include identified viral genome sequences, GenBank and annotation files, AMG summaries, and figures. This is bioconda vibrant version 1.2.1. | Documentation, Uses, and more | |||
| viennarna | lcc | 1.0 | The ViennaRNA Package is a mature suite of programs and C libraries for the prediction and comparison of RNA secondary structure based on a thermodynamic (minimum-free-energy) model and the McCaskill partition-function approach. Its best-known program, RNAfold, predicts the minimum-free-energy structure and base-pairing probabilities of an RNA sequence and can emit dot-bracket notation plus structure and dot-plot PostScript figures; the package also includes RNAcofold (dimer/interaction structures), RNAalifold (consensus structure from an alignment), RNAduplex/RNAup (intermolecular interactions), RNAinverse (design a sequence folding to a target structure), RNAplfold (local structure over long sequences), and a scriptable Python (RNAlib) interface. It supports constraints, temperature settings, and alternative energy parameter sets, and can report the MFE structure, ensemble properties, and centroid/MEA structures for a given sequence. Version 2.4.15 is provided as a conda app in a CentOS 8 + Miniconda bioinformatics/phylogenetics container. | RNA Secondary Structure, Bioinformatics, Computational Biology | Bioinformatics, Biological Sciences | Prediction & Analysis Tools | Documentation, Uses, and more |
| vina | mcc, lcc | 1.0 | Autodock VINA software. | molecular docking, bioinformatics, computational chemistry | Bioinformatics, Molecular Docking | Open-source | Documentation, Uses, and more |
| virsorter | mcc | 1.0 | VirSorter2 identifies viral sequences, including bacteriophages, archaeal viruses, and large/giant eukaryotic viruses (such as NCLDVs), within genomic and metagenomic assemblies. It uses a set of random-forest classifiers trained on multiple viral groups together with a broad reference of viral hallmark and gene-content features, allowing it to recognize diverse viral signals and to detect integrated proviruses within host contigs by boundary prediction. Input is an assembly FASTA; outputs include the extracted viral contigs, per-sequence scores and classifications by viral group, and a boundary/affi table, which are commonly passed downstream to CheckV for quality assessment and to DRAMv for annotation. A run performs viral sorting over selected viral groups, preceded by a one-time setup step that installs its database. Version 2.2.4 is provided as its own conda environment in a Rocky 9 + Miniconda container that also carries RNA/miRNA, QC, and physics-simulation tools (including Geant4/ROOT). | bioinformatics, metagenomics, viral detection, sequence analysis | Genomics, Bioinformatics | Command-line tool | Documentation, Uses, and more |
| visit | lcc | N/A | Visit is an open-source, interactive parallel visualization and graphical analysis tool used for visualizing scientific data. It is particularly focused on large-scale simulations and post-processing of complex data sets, enabling researchers to analyze and present their data in a visually engaging manner. | Visualization, Data Analysis, Scientific Computing | Informatics, Analytics & Information Science, Computer & Information Sciences | Visualization Software | Documentation, Uses, and more |
| vllm | ecc | 1.0 | vLLM (version 0.23.0, this build packaged from the official vllm/vllm-openai:v0.23.0-cu129-ubuntu2404 Docker image for CUDA 12.9) is a high-throughput, memory-efficient inference and serving engine for large language models. Its signature innovation is PagedAttention, which manages the attention key/value cache in non-contiguous pages to minimize memory fragmentation and enable large effective batch sizes; combined with continuous (in-flight) batching, this yields high token throughput under concurrent load. It launches an OpenAI-compatible HTTP server exposing the chat-completions and completions endpoints so existing OpenAI-client code works unchanged, and it supports tensor/pipeline parallelism across multiple GPUs, streaming, quantization, and long-context models. In this deployment it runs on ECC H100 GPUs to back the AI Chat and Launch-an-LLM Open OnDemand apps. It is the serving layer that lets researchers self-host open-weight models with production-grade throughput. | Documentation, Uses, and more | |||
| vmd | lcc | N/A | VMD https://www.ks.uiuc.edu/Research/vmd/ | Molecular Visualization, Molecular Dynamics, Structural Biology | Biophysics, Biochemistry & Molecular Biology, Biological Sciences | Molecular Visualization Tool | Documentation, Uses, and more |
| vpl | mcc, lcc | N/A | VPL (Virtual Programming Lab) is an online platform and learning environment for teaching and practicing programming concepts through interactive coding exercises and assignments. It provides a virtual environment where students can write, compile, and test code, allowing educators to create programming labs and assessments in various programming languages. | Visual Programming, Beginner-Friendly, Graphical Interface | Computer & Information Sciences | Language Programming | Documentation, Uses, and more |
| vscode-tunnel | lcc | N/A | Launch VSCode tunnel sessions on LCC via Singularity | Documentation, Uses, and more | |||
| vtune | mcc, lcc, ecc | N/A | Intel oneAPI VTune Profiler, installed on both MCC and LCC (MCC: /mnt/gpfs3_amd/share/apps/Intel/vtune/<ver>, with versions 2021.1.1, 2021.5.0, 2022.0.0, 2022.4.0, 2024.0 available via 'module spider vtune'). It is a performance-analysis application that profiles CPU/GPU utilization, threading, memory access, microarchitecture hotspots, I/O, and vectorization efficiency in serial and parallel (OpenMP/MPI) applications. The module sets VTUNE_PROFILER_*_DIR and prepends the bin64 directory (whatis: 'Intel(R) oneAPI VTune(TM) Profiler'). It is a genuine user-facing developer/performance-engineering tool, though it is delivered as part of Intel's oneAPI toolkit. | Performance Optimization, Profiling, Intel Architecture | Software Engineering, Computer & Information Sciences | Profiling Tool | Documentation, Uses, and more |
| wget | ecc | N/A | GNU Wget is a free software package for retrieving files using HTTP, HTTPS and FTP, the most widely-used Internet protocols. It is a non-interactive commandline tool, so it may easily be called from scripts, cron jobs, terminals without X-Windows support, etc. | Download Manager, Command-Line Tool | Computer Science, Computer & Information Sciences | System Tool | Documentation, Uses, and more |
| whatshap | lcc | N/A | WhatsHap is a bioinformatics application for read-based phasing of genomic variants. It reconstructs haplotypes from sequencing reads (and optional pedigree information) to determine which alleles lie together on the same chromosome. Its CLI exposes subcommands including phase, stats, compare, hapcut2vcf, unphase, haplotag (tag reads by haplotype), and genotype, making it a standard tool in human and population-genomics variant-phasing workflows. On LCC it is the conda environment whatshap-0.18 (module ccs/conda/whatshap-0.18). | Genomics, Bioinformatics, Variant Phasing | Bioinformatics, Biological Sciences | Tool | Documentation, Uses, and more |
| whipa | lcc | 1.0 | WhIPA is OpenAI Whisper fine-tuned to emit International Phonetic Alphabet transcriptions instead of orthography (Suchardt et al., EMNLP 2025), giving a sequence-to-sequence alternative to frame-level phone recognizers and a useful second opinion alongside ZIPA. This build uses the jshrdt lowhipa-base-cv adapter over the Whisper base checkpoint, both staged read-only outside the image. It takes audio files or directories in any common format and writes a tab separated file with one line per recording holding the IPA string, with Whisper control tokens removed. Because Whisper works on a thirty second window and the adapter was trained on short clips, long recordings should be cut into utterances first, which the diarization and pipeline tools in this container do. | Documentation, Uses, and more | |||
| winnowmap | mcc | 1.0 | Winnowmap is a long-read mapping algorithm derived from minimap2 that is specifically optimized for aligning Oxford Nanopore and PacBio reads to repetitive and previously hard-to-map regions of a reference genome, such as centromeres, segmental duplications, and other high-copy sequence. It replaces the standard uniform minimizer sampling with a weighted minimizer scheme that down-weights (rather than discards) the most frequent k-mers, so that mappings in repeat-rich regions are less biased and more accurate. The workflow has two stages: it first computes a set of highly repetitive k-mers using the bundled meryl k-mer counter, and then runs the mapper with that repetitive-k-mer list to produce SAM output much like minimap2. It was instrumental in telomere-to-telomere (T2T) genome assembly efforts. This is version 2.03 from bioconda. | Mapping, Alignment, Nanopore, Pacbio, Sequencing | Genetics, Biological Sciences | Sequence Alignment Tool | Documentation, Uses, and more |
| wrf | lcc | N/A | The Weather Research and Forecasting (WRF) Model represents a cutting-edge mesoscale system for numerical weather prediction, designed to serve both atmospheric research and practical forecasting needs. Description Source: https://www.mmm.ucar.edu/models/wrf |
Weather Forecasting, Numerical Modeling, Atmospheric Science | Meteorology, Earth & Environmental Sciences | Meteorological Modeling Software | Documentation, Uses, and more |
| wrf-python | lcc | 1.0 | wrf-python (version 1.3.2 from conda-forge) is a Python library of diagnostic and interpolation routines for post-processing output from the Weather Research and Forecasting (WRF-ARW) atmospheric model. It reads WRF NetCDF output and computes derived meteorological variables that are not stored directly, such as CAPE/CIN, sea-level pressure, relative humidity, dewpoint, equivalent potential temperature, storm-relative helicity, and radar reflectivity, and it handles the destaggering of the model's Arakawa-C grid. It also provides interpolation utilities (to pressure/height levels, vertical cross-sections, and 2D surfaces) and integrates with the geoscience Python stack, returning results as xarray DataArrays with coordinate metadata and supporting Cartopy/matplotlib and PyNGL for plotting. Its central getvar interface retrieves and computes named diagnostic fields from a WRF output file. It is the modern Python replacement for the older NCL WRF diagnostics and is standard in mesoscale/atmospheric research and forecasting analysis. | Meteorology, Data Analysis, Visualization, Python, Scientific Computing | Meteorology, Weather Forecasting | Library | Documentation, Uses, and more |
| xcb-proto | mcc, lcc | N/A | The X protocol C-language Binding (XCB) is a replacement for Xlib featuring a small footprint, latency hiding, direct access to the protocol, improved threading support, and extensibility | X11, XCB, protocol, graphics, development | Development Library | Documentation, Uses, and more | |
| xenium ranger | mcc | 1.0 | Xenium Ranger is 10x Genomics' command-line pipeline for post-processing, re-analyzing, and customizing the output of Xenium in-situ single-cell spatial imaging experiments after the instrument's onboard analysis. It lets researchers re-run steps of the Xenium Onboard Analysis without repeating the physical run: re-segmenting cells (using nucleus expansion, custom stains, or imported segmentation masks), importing external/custom cell-segmentation boundaries, relabeling or importing custom gene panels, and resampling/recomputing transcript-to-cell assignments. Its subcommands cover importing segmentation, resegmenting, relabeling, and importing panels, each taking a Xenium output bundle (the directory containing the morphology images, transcripts, and cell feature matrix) and producing an updated bundle with a new cell-feature matrix and boundaries. The refreshed outputs load into Xenium Explorer or into scanpy/squidpy/Seurat for downstream spatial single-cell analysis. This deployment is version 4.0.0, staged from a local 10x tarball. | Documentation, Uses, and more | |||
| xextproto | mcc, lcc | N/A | X11 Extension protocols and auxiliary headers | Linux, X11, X Window System | Computer & Information Sciences | Library | Documentation, Uses, and more |
| xf86vidmodeproto | mcc | N/A | Documentation, Uses, and more | ||||
| xineramaproto | mcc | N/A | Documentation, Uses, and more | ||||
| xkbcomp | mcc | N/A | xkbcomp is a utility for compiling XKB (X Keyboard Extension) descriptions into binary files that can be used by the X server. | X Window System, Keyboard Configuration, Linux, Unix | Software Engineering, Other Computer and Information Sciences | Utility | Documentation, Uses, and more |
| xkbdata | mcc | N/A | Documentation, Uses, and more | ||||
| xproto | mcc, lcc | N/A | X protocol and ancillary headers | Software Development, Protocol Development, X Window System | Computer & Information Sciences | Library | Documentation, Uses, and more |
| xrandr | mcc, lcc | N/A | xrandr is a command-line tool for managing and configuring display settings in the X Window System. It allows users to set the size, orientation, and reflection of the outputs for the display. | Display Management, X Window System, Linux, Command Line Tool | Software Engineering, Other Computer and Information Sciences | Command Line Utility | Documentation, Uses, and more |
| xtea | mcc | 1.0 | xTea (version 0.1.9 from bioconda) detects non-reference transposable-element insertions, targeting the major active human mobile-element families L1 (LINE-1), Alu, SVA, and HERV, from whole-genome sequencing data. It supports Illumina short reads, PacBio/ONT long reads, and 10x Genomics linked reads, using discordant/clipped read signatures (and, for long reads, split-read evidence) to localize and characterize insertions including their transductions and truncation status; a bundled deep-forest classifier (installed via pip) refines calls. The container also carries its aligner/utility dependencies (bwa, samtools, minimap2). A typical run supplies a BAM/CRAM, a sample list, a reference genome, and family-specific repeat libraries through the xtea driver scripts, producing VCF/BED-style call sets per element type. It is used in human genomics and disease studies where somatic or germline retrotransposon insertions are of interest. | Documentation, Uses, and more | |||
| xtrans | mcc, lcc | N/A | xtrans is a general-purpose transducer-based approach for defining translations between various formats and data representations. It provides a flexible framework for specifying bidirectional transformations between different data structures. | Data Transformation, Format Conversion, Transducer-Based Tool | Computer & Information Sciences | Tool | Documentation, Uses, and more |
| xxhash | mcc | N/A | xxHash is an extremely fast non-cryptographic hash algorithm, providing high-speed hashing for data integrity checks. | hashing, performance, data integrity | Software Engineering, Other Computer and Information Sciences | Library | Documentation, Uses, and more |
| xz | mcc, lcc, ecc | N/A | XZ is an open-source data compression utility that uses the LZMA (Lempel-Ziv-Markov chain algorithm) compression algorithm. It is commonly used to compress files and reduce their size while maintaining high compression ratios and providing options for efficient data archiving and distribution. | Compression, Data Compression, Open-Source | Computer Science, Computer & Information Sciences | Utility | Documentation, Uses, and more |
| yahs | mcc | 1.0 | YaHS (Yet another Hi-C scaffolding tool) is a fast, memory-efficient program for scaffolding draft genome assemblies into chromosome-scale sequences using Hi-C proximity-ligation data. It takes a contig-level assembly FASTA (with its .fai index) and a coordinate- or position-sorted BAM/BED/PA5/BIN file of Hi-C read alignments to those contigs, and iteratively joins and orders/orients contigs based on the Hi-C contact signal. A key feature is that YaHS also breaks misassembled contigs at points of low Hi-C support, improving correctness relative to scaffolding alone. It produces a scaffolds FASTA plus an .agp file describing the contig-to-scaffold layout and .bin files that can be converted (via the bundled juicer_tools helper and juicer_pre) into a Juicebox-compatible .hic contact map for manual review and correction. It is frequently paired with tools like Arima/BWA Hi-C mapping pipelines upstream and Juicebox/PretextMap downstream. This is version 1.2.2 from Bioconda. | Hi-C scaffolding, genome assembly, chromosome-scale assembly | Life Sciences, Bioinformatics | Documentation, Uses, and more | |
| yak | mcc | 1.0 | yak (yet another k-mer analyzer) is a lightweight k-mer toolkit from Heng Li's group, most commonly used to evaluate genome-assembly quality independent of a reference. Version 0.1. It builds Bloom-filtered k-mer hash tables from accurate short reads (typically Illumina) and uses them to estimate assembly base accuracy as a Phred-scaled Quality Value (QV) and to assess k-mer completeness. In trio settings it also computes hap-mer sets from parental reads to measure haplotype phasing/switch errors of a diploid assembly. The core workflow first counts k-mers from the short reads into a yak hash table and then evaluates an assembly against that table, reporting QV and coverage-adjusted accuracy. Its hap-mer output is a key input for trio-binning quality control in assemblers like hifiasm. | Documentation, Uses, and more | |||
| z3 | mcc | N/A | Z3 is a high-performance theorem prover developed by Microsoft Research that is used for solving satisfiability modulo theories (SMT) problems. It provides capabilities for automated reasoning and constraint solving in various domains such as software verification, automated planning, and formal methods. | Theorem Prover, Satisfiability Modulo Theories (Smt) Solver, Software Verification, Hardware Verification, Security, Ai, Machine Learning | Artificial Intelligence, Computer & Information Sciences | Theorem Prover | Documentation, Uses, and more |
| zfp | mcc, lcc | N/A | zfp is a compressed numerical array library providing high-throughput, low-overhead fixed-rate, fixed-precision encoding and compression of 1D and 2D scientific data. | Numerical Array, Compression, Scientific Data | Computer Science, Computer & Information Sciences | Library | Documentation, Uses, and more |
| zip | mcc | N/A | Zip is a popular utility for compressing and archiving files and directories into a single zip file. It provides a convenient way to reduce file sizes, organize data for storage or transmission, and create compressed archives that can be easily shared and extracted. | Compression, File Packaging, Utility | Computer Science, Other Computer & Information Sciences | Utility | Documentation, Uses, and more |
| zipa | lcc | 1.0 | ZIPA is a multilingual speech-to-IPA phone recognizer (Zhu et al., ACL 2025) that transcribes speech directly into International Phonetic Alphabet symbols without needing a language model or an orthographic transcript, which makes it suited to fieldwork recordings, under-resourced languages and phonetic research. This build runs the anyspeech zipa-small-crctc-ns-700k CTC model, trained on 88 languages, through ONNX Runtime on the GPU. It reads audio files or whole directories in any common format, converts them to 16 kHz mono internally, and writes a tab separated file with one line per recording holding the space separated phone sequence, where a block character marks word boundaries. Half and eighth precision copies of the model are bundled for faster runs, and the weights are shared read-only outside the image so no download is needed at run time. | Documentation, Uses, and more | |||
| zlib | mcc, lcc | N/A | zlib is a software library for data compression and decompression using the DEFLATE data compression algorithm. It offers a fast and efficient compression method for reducing file sizes and conserving storage space. zlib is widely used in various applications for compression and decompression of data streams and files. | Compression, Data Compression, File Compression, Deflate Algorithm | Software Engineering, Computer & Information Sciences | Data Compression | Documentation, Uses, and more |
| zlib-ng | ecc | N/A | zlib-ng is a fast, efficient, and portable data compression library that is a drop-in replacement for zlib. It is designed to be compatible with zlib while providing improved performance and additional features. | compression, data processing, library, performance | Computer Science, Software Engineering | Library | Documentation, Uses, and more |
| zsh | mcc | N/A | Z shell (zsh) version 5.9, built with ncurses from source. | Unix Shell, Customizable Shell, Command Line Interface | Software Engineering, Computer & Information Sciences | Command Line Interface | Documentation, Uses, and more |
| zstd | mcc, lcc, ecc | N/A | zstd (Zstandard) is a high-performance data compression library and command-line tool that offers fast compression and decompression speeds with the ability to achieve high compression ratios. It is designed to balance efficient compression and decompression with speed, making it suitable for various data processing and storage applications. | Compression, Data Compression, Algorithm | Computer Science, Computer & Information Sciences | Compression Algorithm | Documentation, Uses, and more |