We demonstrate that membership inference attacks against fine-tuned large language models achieve 0.95 AUC using only output token probabilities, without access to model parameters or gradients.
Diffusion models have achieved state-of-the-art image generation quality as measured by FID and IS scores. However, we demonstrate that these metrics mask a critical failure mode: anatomically implausible human hands.
Distributed tracing is foundational to microservice observability, yet its performance overhead is poorly quantified, particularly at tail latencies. We instrument 23 production microservice deployments across 4 organizations, measuring tracing overhead at the 50th, 95th, and 99th percentiles of CPU utilization.
Continual learning methods are universally evaluated under a discrete task-boundary assumption, where distribution shifts occur instantaneously between clearly delineated tasks. We argue this assumption is ecologically invalid and demonstrate that five leading continual learning methods (EWC, SI, PackNet, ER, DER++) fail catastrophically when task boundaries are gradual.
We present new results on ramsey theory with applications to sat solvers. Our main theorem establishes sharp bounds that improve upon the best previously known results, settling a conjecture in the affirmative for the cases considered.
We present new results on graph reconstruction with applications to reconstruction conjecture. Our main theorem establishes sharp bounds that improve upon the best previously known results, settling a conjecture in the affirmative for the cases considered.
We conduct the largest study to date on coreference, analyzing 38,271 instances across 17 datasets spanning multiple domains. Our key finding is that clinical nlp accounts for 17.
We empirically characterize how inference-time compute scales with task performance for agentic AI workloads. Across 14 agentic benchmarks spanning web navigation, code generation with tool use, and multi-step reasoning, we find that performance follows a power law with exponent 0.
Foundation models for zero-shot object detection, including CLIP-based detectors and Grounding DINO, have achieved remarkable performance on natural image benchmarks. However, their deployment in industrial quality inspection remains largely untested.
We present new results on oriented coloring with applications to planar graphs. Our main theorem establishes sharp bounds that improve upon the best previously known results, settling a conjecture in the affirmative for the cases considered.
We present new results on chromatic polynomials with applications to graph isomorphism. Our main theorem establishes sharp bounds that improve upon the best previously known results, settling a conjecture in the affirmative for the cases considered.
We conduct the largest study to date on simplification, analyzing 43,266 instances across 7 datasets spanning multiple domains. Our key finding is that ambiguity accounts for 24.
This paper investigates the relationship between prompt injection and rag through controlled experiments on 28 diverse datasets totaling 19,998 samples. We propose a novel methodology that achieves 8.
We present a systematic empirical study examining neural architecture search across 13 benchmarks and 13,585 evaluation instances. Our analysis reveals that skip connections plays a more critical role than previously recognized, achieving 0.
We conduct the largest study to date on sim to real, analyzing 14,968 instances across 18 datasets spanning multiple domains. Our key finding is that manipulation accounts for 5.
This paper investigates the relationship between 3d reconstruction and normal maps through controlled experiments on 18 diverse datasets totaling 31,631 samples. We propose a novel methodology that achieves 31.
We present a systematic empirical study examining deformable objects across 5 benchmarks and 28,196 evaluation instances. Our analysis reveals that force torque plays a more critical role than previously recognized, achieving 0.
We conduct the largest study to date on code review, analyzing 24,005 instances across 12 datasets spanning multiple domains. Our key finding is that llm accounts for 14.
We present a rigorous experimental and theoretical investigation addressing the claim embedded in this work's title. Using a combination of analytical derivations, numerical simulations, and where applicable, experimental data from state-of-the-art quantum hardware, we establish precise quantitative thresholds and scaling behaviors.
This paper investigates the relationship between morphology and pretraining through controlled experiments on 23 diverse datasets totaling 26,178 samples. We propose a novel methodology that achieves 9.