Named after the ACTN3 "Speed Gene" - We accelerate your discoveries
Website Portfolio LinkedIn Email
𧬠AI-Powered Genomics β’ π¦ Production Pipelines β’ π¬ Reproducible Research β’ π Open Source
Transforming raw sequencing data into clinical insights through modern bioinformatics workflows
The ACTN3 gene (the "Speed Gene") is associated with explosive muscle performance and athletic prowess. Just like this gene enables peak physical performance, ACTN3 Bioinformatics delivers fast, efficient, and high-performance genomic analysis solutions.
Accelerate drug discovery and precision medicine by providing:
- β‘ Rapid turnaround - Days, not weeks
- π€ AI-powered workflows - LLM-assisted analysis, automated QC
- π¬ Reproducible pipelines - Snakemake, Docker, CI/CD
- π Publication-ready - Quarto reports, GitHub repos, full documentation
- β Pharma-grade quality - 10+ years Biotech/Pharma experience
- π 6 publications in top-tier journals (iScience, Clinical Cancer Research)
- π’ Pharma expertise - Roche, Novartis standards
- π€ Open source first - Contributing to pharmaverse
- π R/Pharma 2025 - 11 workshop certifications
- π Cutting-edge tech - AI/LLM, Positron IDE, Quarto
A reference architecture template for nationwide population whole-genome sequencing programs
A comprehensive, end-to-end bioinformatics architecture template for national population WGS programs at the 5β10 K cohort scale (G4PL-class projects). Synthesizes lessons from gnomAD v4, UK Biobank, FinnGen, Genomic Medicine Sweden, HPRC, the GDI Starter Kit, and the GA4GH standards stack into a deployable reference design β published as an independent educational resource for research groups and consortia.
Key features:
- 𧬠Multi-platform integration: Illumina short-read · PacBio HiFi · Oxford Nanopore · Bionano OGM
- π¬ Six-channel per-sample variant calling: SNV/indel Β· SV Β· repeats Β· HLA Β· PGx Β· mtDNA
- π Cohort-scale joint analysis: GLnexus / Hail VDS Β· SHAPEIT5 phasing Β· gnomAD-aligned QC
- π Population genetics: PCA Β· ADMIXTURE Β· FST Β· IBD Β· LD Β· ROH Β· haplogroups
- π GA4GH-compliant federated access: Beacon v2 Β· htsget Β· DRS Β· Crypt4GH Β· DUO
- π Seven phased documentation modules + comprehensive Phase 1 architecture document
Technologies: Nextflow Β· nf-core/sarek Β· Hail VDS Β· DeepVariant Β· GLnexus Β· GATK-SV Β· Sniffles2 Β· SHAPEIT5 Β· Truvari Β· VEP Β· GDI Starter Kit Β· Phenopackets v2 Β· VRS v2 Β· Crypt4GH
Production-grade Snakemake pipeline transforming scRNA-seq CRISPR screens into balanced, harmonized datasets for AI/ML training
Key Features:
- π Automated QC + normalization (Scanpy)
- βοΈ Smart class balancing for unbiased ML
- π Batch integration (Harmony/BBKNN)
- 𧬠Feature engineering (pathways, TF regulons)
- π Leave-genes-out cross-validation
- π Python 3.10, Snakemake, Quarto
Performance: Processes 10k cells in ~15 minutes on laptop (AMD Ryzen 5 7535HS, 16GB RAM)
Impact: Virtual Cell Challenge 2025 submission showcasing production-ready bioinformatics engineering
Comprehensive Quarto knowledge base documenting R/Pharma 2025 conference workshops, trends, and best practices
Content Highlights:
- π 19 workshop summaries (AI/LLM, Clinical Reporting, Validation, Bayesian Methods)
- π€ 30+ presentation summaries
- π Industry trend analysis (AI revolution, GSK's 50%+ R adoption)
- π οΈ Complete A-Z tools catalog (ellmer, gtsummary, teal, officer, etc.)
- πΌ Career insights for R pharma professionals
Technologies: Quarto, RMarkdown, GitHub Actions, GitHub Pages
Recognition: Showcases expertise in modern documentation workflows and pharmaceutical R ecosystem
From FASTQ to Figures
- RNA-seq (bulk & single-cell)
- ChIP-seq, ATAC-seq, CUT&RUN
- CRISPR/ORF screens
- WGS/WES variant calling
- scTCR/scBCR-seq
- Spatial transcriptomics
- Long-read (PacBio, Nanopore)
Format: Quarto reports + GitHub repo
Production-Ready Workflows
- Snakemake/Nextflow pipelines
- Docker containerization
- HPC optimization
- CI/CD integration (GitHub Actions)
- Automated testing & validation
- Comprehensive documentation
Format: Reproducible code + pkgdown site
Modern AI-Powered Analysis
- LLM-assisted code generation
- Automated QC systems
- ML-ready dataset preparation
- Feature engineering
- Multi-omics integration
- Prompt engineering workflows
Format: Trained models + deployment scripts
Bioconductor Standards
- Custom package design
- roxygen2 documentation
- testthat unit testing
- pkgdown websites
- CRAN/Bioconductor submission
- Validation documentation
Format: Fully documented package
Shiny Applications
- Real-time QC monitoring
- Clinical data exploration
- Drug response visualizations
- Pathway enrichment browsers
- GitHub Pages portfolios
- Quarto knowledge bases
Format: Hosted app + source code
From Analysis to Paper
- Methods sections
- Supplementary analyses
- Publication-quality figures
- Statistical consulting
- Reproducible workflows
- Data deposition (GEO/EGA)
Format: Camera-ready materials
R Python Snakemake Docker Git Quarto
Bioconductor Seurat Scanpy scikit-learn
R/Bioconductor: limma, edgeR, DESeq2, fgsea, ComplexHeatmap, Seurat, crisprVerse, MAGeCK
Python: Scanpy, AnnData, PyTorch, TensorFlow, scikit-learn, Harmony, gseapy
Pipelines: Snakemake β₯7.0, Nextflow DSL2, Bash scripting
Reproducibility: Quarto, RMarkdown, knitr, pkgdown, GitHub Actions
Infrastructure: Docker, Conda/Mamba, HPC (Slurm), AWS/Azure
AI/LLM: ellmer, OpenAI API, Anthropic Claude, Prompt Engineering
1. Transcription factor Zfx regulates tumor immune evasion | iScience (2025)
Role: CRISPR screen analysis (~160K sgRNAs), ChIP-seq, TCGA survival models
DOI
2. Transcriptional subtypes in lung adenocarcinoma | Clinical Cancer Research (2021)
Role: NMF subtype discovery, 113-gene PAM classifier (87-91% accuracy)
DOI Code
3. T cell-dependent bispecific therapy | Cancer Immunology Research (2024)
Role: NK cell RNA-seq, GSEA, pathway enrichment
DOI
π View All Publications on ORCID
We stay at the forefront of pharmaceutical bioinformatics trends:
Status: Production-ready
Our Capability:
- β LLM-assisted code generation (ellmer, ChatGPT API)
- β Automated QC report generation
- β Prompt engineering workflows
- β Privacy-preserving implementations
Industry Evidence:
- 8+ R/Pharma sessions on AI
- GSK, Roche, Merck deploying AI systems
- Regulatory acceptance growing
Status: Mainstream
Our Contribution:
- β Public GitHub repositories (VCC, R/Pharma Portal)
- β Pharmaverse-aligned workflows
- β Reproducible research standards
- β Community engagement
Industry Milestone:
- GSK achieved 50%+ R code adoption
- Pharmaverse thriving
- SAS β R migration accelerating
Status: Validated
Our Compliance:
- β 9+ years Roche/Genentech experience
- β ISO-compliant documentation
- β GxP-ready workflows
- β Reproducible audit trails
Regulatory Landscape:
- FDA software neutrality established
- R-based submissions routine
- Validation frameworks standardized
Status: Production
Our Solutions:
- β Snakemake/Nextflow pipelines
- β CI/CD integration (GitHub Actions)
- β Automated testing & validation
- β Template-based reporting (Quarto)
Industry Impact:
- CDISC ARS/ARM driving automation
- TFL turnaround: weeks β days
- AI-powered table generation
Szymon Myrta
Founder & CEO
10+ Years in Pharma/Biotech
- π’ Current: Bioinformatician
- π Publications: 6 peer-reviewed papers in iScience, Clinical Cancer Research, EMBO MM
- π Education: Double MSc - MSc Bioinformatics (Silesian University), MSc Molecular Medicine (Cranfield University)
Core Competencies:
- NGS Analysis (RNA-seq, scRNA-seq, ChIP-seq, CRISPR screens, WGS/WES)
- AI/ML Integration (LLM APIs, scikit-learn, TensorFlow)
- Pipeline Development (Snakemake, Nextflow, Docker, HPC)
- R Package Development (Bioconductor standards, testing, documentation)
- Clinical Bioinformatics (biomarker discovery, patient stratification)
Languages: English (Fluent), Polish (Native)
π€ AI in Bioinformatics: - LLM-assisted code generation & debugging - AI-powered variant annotation - Automated literature mining for pathway enrichment - Prompt engineering best practices 𧬠Multi-Modal Omics: - Spatial transcriptomics (10X Visium, Xenium) - Single-cell + proteomics + epigenomics integration - Multi-omic predictive modeling for precision medicine π§ͺ Non-coding RNA: - lncRNA, miRNA, circRNA in cancer & development - Integration and correlation analysis (ncRNA β mRNA) - Regulatory network reconstruction βοΈ Software Development: - R package development (Bioconductor standards) - Positron IDE + Quarto for literate programming - Nextflow DSL2 + Wave containers for cloud scalability - Git-based collaboration for reproducible science π― CRISPR & Functional Genomics: - Perturb-seq analysis pipelines - Base/prime editor screen optimization - Screen hit validation workflows π₯ Precision Medicine: - Real-time clinical decision support systems - Pharmacogenomics + liquid biopsy integration - Patient-specific therapy modeling
| Advantage | Description |
|---|---|
| β‘ Speed | Established pipelines ensure rapid turnaround (RNA-seq: 2-4 weeks, scRNA-seq: 4-8 weeks) |
| π Quality | Pharma-grade standards from 9+ years at Roche/Genentech |
| π€ Modern | AI/LLM integration, Positron IDE, Quarto documentation, cutting-edge methods |
| π Proven | 6 publications in top journals, 800+ patients analyzed, 160K sgRNAs processed |
| π¬ Reproducible | Full GitHub repos, Docker containers, CI/CD, comprehensive documentation |
| β Compliant | ISO experience, GxP-ready workflows, validated pipelines, audit trails |
| π Remote | 9+ years remote work experience, international teams, flexible time zones |
| π€ Collaborative | Contributing to pharmaverse, open-source first, knowledge sharing |
We're available for:
- 𧬠Contract NGS analysis projects (4-20 weeks)
- π¦ Custom R package development (4-12 weeks)
- βοΈ Production pipeline engineering (4-12 weeks)
- π€ AI/ML integration consulting (6-16 weeks)
- π Interactive Shiny dashboard creation (4-6 weeks)
- π Publication support (methods, figures, analyses)
- π Training & workshops (R/Bioconductor, reproducible research)
Company Email: kontakt@actn3.pl
Personal Email: szymon.myrta@gmail.com
Response Time: Within 24 hours
ACTN3 Bioinformatics
β‘ ACTN3 Bioinformatics - Where Speed Meets Science
"Transforming genomic data into insights through reproducible pipelines and AI-powered workflows"
All public repositories are licensed under the MIT License unless otherwise specified.
β If you find our work valuable, please star our repositories!