Insilico Medicine launched what it describes as a benchmark-as-a-service to test whether frontier AI models make meaningful decisions in drug discovery beyond memorizing training data. The new Drug Discovery and Development Benchmark-as-a-Service targets weaknesses in evaluation: many benchmarks measure recall of known answers rather than the scientific reasoning and experimental decision-making required for drug development. The company frames the tool as a standardized framework for stress-testing AI systems under conditions that better reflect real discovery workflows, including development-relevant decision points. The move signals a push toward more defensible validation metrics as more AI vendors enter biopharma R&D. For biotech teams assessing AI capabilities, the key development is the introduction of an externalized evaluation service intended to make model performance comparable and to separate usable innovation from dataset learning.