Insilico Medicine launched what it calls the drug discovery and development benchmark-as-a-service for frontier AI models, designed to test whether AI systems make meaningful scientific decisions rather than relying on training-set recall. The company’s framework is intended as a real-world “blind test,” with a public leaderboard meant to improve how models are evaluated. The move arrives as AI vendors increasingly pitch foundation-model capability, while buyers and the scientific community demand stronger, more decision-focused validation.