An early proof-of-concept study using CT radiomics to predict response to radiotherapy in elderly esophageal cancer patients showed unstable performance when fully nested validation was used. The results underscore a recurring issue in medical AI: high initial metrics can fail to generalize once leakage controls and more rigorous validation protocols are applied, limiting confidence in any claimed benefit from adding imaging models. For radiomics developers and oncology trialists, the study is a reminder to treat internal validation as insufficient and to design prospective, externally validated studies before translating CT-based response predictors into routine workflows.