Learned cost models for parallel stream processing typically rely on hand-engineered features, which requires substantial domain expertise and manual effort. This paper presents an automated feature selection pipeline that identifies the features most relevant to predicting query performance in parallel stream processing, reducing the manual effort involved in building learned cost models without sacrificing prediction accuracy.