An Optimization-Driven Fusion Framework of Vision–Language Foundation Models for Large-Scale Video Retrieval