We present battisiBot v2, a 24-step sequential reinforcement learning environment for automated orthodontic aligner trajectory planning. An agent plans one aligner stage at a time across 28 teeth as SE(3) poses, with 5 tool-use actions, Andrews Six Keys occlusion scoring, PDL biomechanical model, collision detection, adversarial non-compliance, 8-axis adaptive difficulty, 8 malocclusion classes, 5 arch forms, and real clinical data from Open-Full-Jaw (17 patients) and Mendeley Jaw Models.
**Background:** Ophthalmic drug safety surveillance faces a fundamental challenge: the same drug can exhibit radically different adverse event (AE) profiles depending on the clinical indication, route of administration, and patient population. Traditional pharmacovigilance methods, which aggregate adverse events across all uses of a drug, systematically mask indication-specific toxicity signals.
We present MTX-LIVER, an executable Python skill for transparent liver-safety risk stratification before or during low-dose methotrexate therapy in rheumatic and autoimmune disease. The model integrates obesity, diabetes, known steatosis/NAFLD, alcohol exposure, chronic hepatitis B/C, baseline and current aminotransferases, albumin, platelet count, methotrexate weekly dose, treatment duration, cumulative dose, folate supplementation, concomitant leflunomide, and persistent transaminitis.
**Background**: Hepatocellular carcinoma (HCC) is the sixth most common cancer globally, with over 870,000 new cases annually. Targeted therapies and immune checkpoint inhibitors have transformed HCC treatment, yet these drugs carry inherent hepatotoxicity risks that are amplified in patients with compromised liver function.
A persistent reproducibility crisis in biomedical research has been attributed to statistical errors, selective reporting, and p-hacking—yet a comparatively underexplored mechanism is the role of unstated assumptions that silently link evidence to conclusions. When a paper's core claims rest on premises that are never made explicit, the validity of those claims depends entirely on the truth of assumptions that are never tested, discussed, or even acknowledged.
Biomedical researchers spend a disproportionate amount of time navigating fragmented literature to identify viable therapeutic hypotheses. We introduce BioLit-Scout, a modular, agent-executable skill that automates the aggregation, filtering, and synthesis of published evidence for hypothesis prioritization in disease mechanism research.
Reliable biomedical language modeling requires not only factual recall but also robust handling of invalid evidence. We present a bioinformatics-oriented contamination benchmark that measures whether LLMs rely on retracted medical papers under clinically framed tasks, using a versioned Kaggle dataset snapshot and a two-stage evaluation protocol.
Biological literature synthesis for therapeutic target identification remains a manual, time-consuming process with limited reproducibility. Researchers navigating thousands of publications across PubMed, bioRxiv, and domain databases face fragmented evidence, inconsistent nomenclature, and difficulty prioritizing candidate targets.
Janus kinase inhibitors are effective therapies for rheumatoid arthritis and other autoimmune diseases, but thrombotic safety concerns remain clinically important. We present VTE-JAK, an executable Python skill for transparent pre-treatment and treatment-review stratification of venous thromboembolism risk in patients being considered for JAK inhibitor therapy.
Trojan Paper Medical Benchmark presents a web-first workflow for evaluating LLM metacognitive robustness against retracted medical evidence. It discovers retracted studies from public online sources, constructs benchmark cases with unreliable-claim and retraction context, and runs a two-stage target-plus-judge evaluation pipeline with contamination-sensitive metrics.
Resumption of oral anticoagulation (OAC) after a major gastrointestinal bleed (GIB) in atrial fibrillation (AF) is a recurring clinical question without a published, transparent, domain-weighted net-benefit tool. Observational cohorts consistently report lower all-cause mortality and lower thromboembolic events in patients restarted on OAC versus permanently withheld, but also elevated rebleed rates with hazard ratios clustering between 1.
Rechallenge with immune checkpoint inhibitors (ICIs) after a grade 3 or higher immune-related hepatitis (irHepatitis) is a recurring clinical question without a published, transparent, domain-weighted risk tool. Published retrospective series report pooled recurrence rates of any-grade immune-related adverse event (irAE) on rechallenge in the 25-55% range, with recurrence of the same-organ irAE clustered at the upper end, but effect sizes for individual modifiers (time-to-resolution, peak ALT, steroid taper duration, combination vs.
Executable clinical skill for steroid-induced hyperglycemia risk stratification using baseline glycemic vulnerability, glucocorticoid exposure burden, and host susceptibility in rheumatic and autoimmune disease.
Tumour-associated neutrophils (TANs) in hepatocellular carcinoma (HCC) span a continuous activation spectrum from anti-tumour antigen-presenting states to pro-tumour angiogenic and immunosuppressive states [Grieshaber-Bouyer et al., Nature Communications, 2021; Antuamwine et al.
Clinical enzyme testing is one of the most frequently ordered laboratory panels in healthcare, yet its interpretation remains heavily dependent on physician experience and implicit knowledge. We present **ClinicalEnzymeDiagnostics-Skill**, an open-source AI agent that transforms routine clinical chemistry data into structured differential diagnoses using Bayesian probabilistic reasoning.
GWASEngine is a complete GWAS analysis pipeline implemented entirely in Python using NumPy, SciPy, and scikit-learn. Six modules: QC, linear regression GWAS, LD clumping, polygenic risk scores (C+T), Bayesian fine-mapping (Wakefield ABF), and LD Score Regression.
Heart rate variability (HRV) metrics are widely used in clinical cardiology, psychophysiology, and consumer wellness applications, yet the accuracy of these metrics relative to known autonomic ground truth has never been systematically quantified. This study presents the first comprehensive benchmark of 14 standard HRV metrics — 7 time-domain and 7 frequency-domain — computed from synthetic RR-interval series with exactly specified autonomic parameters.
The Hallmarks of Aging framework identifies twelve interdependent biological processes that drive organismal decline. While individual longevity compounds have been extensively profiled, the combinatorial question -- which minimal set of compounds maximally covers the hallmark landscape -- remains unaddressed.
Skull base surgery demands precise corridor selection to maximize lesion exposure while minimizing cranial nerve injury. Despite decades of refinement, approach selection remains guided primarily by individual expertise rather than formal quantitative frameworks.