Insilico Medicine launched a Drug Discovery and Development Benchmark as a Service aimed at stress-testing frontier AI models on real-world scientific decision-making. The company described a benchmark framework that evaluates whether models can perform meaningful drug-discovery tasks rather than relying on memorization. Insilico says it built the platform using a large set of validated pharmaceutical compound-to-phenotype (PCC) programs and its own model-development experience, packaging the evaluation as a “blind test” with a transparent public leaderboard. The initiative targets a known weakness in AI evaluation: benchmarks that don’t reflect the operational constraints of real scientific workflows. For biotech and pharma groups, the service is positioned as a tooling layer for selecting models that can integrate into discovery pipelines. If widely adopted, it could influence vendor selection and shift evaluation from generic NLP-style metrics toward task-based scientific performance.