Accurate simulation of precipitation remains a major challenge in mountainous regions, where complex topography, mixed precipitation types, and strong seasonal contrasts constrain climate model performance. This study evaluates bias-corrected Coupled Model Intercomparison Project Phase 6 (CMIP6) General Circulation Models (GCMs) at station scale across the Jhelum River Basin (JRB) (1985–2014). Thirteen GCMs were assessed annually and seasonally at six stations using correlation coefficient (CC), root mean square error (RMSE), normalized RMSE (NRMSE), absolute normalized mean bias error (ANMBE), Nash–Sutcliffe efficiency (NSE), Kling-Gupta efficiency (KGE), and probability skill score (PSS). To ensure objective, reproducible model selection, we developed an integrated framework combining three independent ranking methods Technique for Order Preference by Similarity to Ideal Solution (TOPSIS), PROMETHEE-II, and VIseKriterijumska Optimizacija I Kompromisno Resenje (VIKOR) with explicit rank-stability quantification across methods, stations, and seasons, and a sensitivity analysis against alternative weighting and Monte-Carlo weight uncertainty. This combination of multi-method ranking with formal stability quantification is a central methodological contribution of this study. Results show bias correction substantially reduces, but does not eliminate, systematic error: annual RMSE ranges from 9.1–14.3 mm and ANMBE from 1.5–35.2%, while correlation and efficiency metrics remain weak throughout (CC: −0.02 to 0.04; NSE: −1.13 to −0.42), consistent with the expected behavior of free-running GCMs at daily resolution rather than a basin-specific deficiency. Elevation modulates performance non-uniformly: NRMSE and NSE improve with elevation while ANMBE and KGE worsen, and the highest-elevation station records both the best NRMSE/NSE and the largest annual bias in the basin. Autumn is consistently the weakest season, reflecting monsoon-withdrawal transition dynamics, localized convection, and orographic effects that coarse-resolution GCMs and a sparse station network cannot fully resolve. The integrated consensus ranking identifies MPI-ESM1-2-LR as the most robust model and MRI-ESM2-0 as the least robust, both stable across 21 of 22 alternative configurations tested; the remaining 11 models’ ranking is not robust to reasonable weighting variation and should not be treated as a strict ordering. Overall, the study establishes a reproducible, rank-stability-aware framework for station-scale climate model evaluation, providing actionable guidance for precipitation-driven hydrological modeling and climate-risk assessment in Himalayan terrain.