# ML Pipelines และการตรวจสอบ (Data Science & ML) > Scikit-learn pipelines, cross-validation, GridSearchCV, RandomizedSearchCV, data leakage, การแบ่งชั้น - 22 คำถามสัมภาษณ์ - Mid-Level - [คำถามสัมภาษณ์: Data Science & ML](https://sharpskill.dev/th/technologies/data-science/interview-questions.md) ## 1. ข้อได้เปรียบหลักของการใช้ Pipeline ของ scikit-learn แทนการใช้การแปลงด้วยตนเองคืออะไร? **คำตอบ** Pipeline รับประกันว่าการแปลงเดียวกันจะถูกนำไปใช้อย่างสม่ำเสมอกับทั้งข้อมูล training และ test มันรวมขั้นตอน preprocessing และ modeling ทั้งหมดไว้ในออบเจ็กต์เดียว ซึ่งทำให้โค้ดง่ายขึ้น ป้องกัน data leakage และทำให้การ deploy model ไปยัง production ง่ายขึ้น ## 2. ควรเรียก method ใดบน Pipeline เพื่อ train ทุกขั้นตอนและทำการทำนาย? **คำตอบ** method fit_predict ไม่มีอยู่ใน Pipeline สำหรับ regression หรือ classification คุณต้องเรียก fit() ก่อนเพื่อ train pipeline จากนั้นเรียก predict() เพื่อรับการทำนาย หรืออีกทางหนึ่ง สามารถเรียก fit() ตามด้วย predict() แยกกันเพื่อการควบคุมที่มากขึ้น ## 3. Data leakage ในบริบทของ machine learning คืออะไร? **คำตอบ** Data leakage เกิดขึ้นเมื่อข้อมูลจาก test set หรือข้อมูลในอนาคตถูกใช้โดยไม่ตั้งใจระหว่างการ training สิ่งนี้สามารถเกิดขึ้นระหว่าง preprocessing (การคำนวณค่าเฉลี่ยจากทั้ง dataset ก่อนการ split) หรือผ่าน feature ที่มี target ทางอ้อม ผลลัพธ์คือประสิทธิภาพสูงเทียมที่ไม่สามารถ generalize ได้ ## มีอีก 19 คำถาม - บทบาทของ ColumnTransformer ใน scikit-learn คืออะไร? - K-Fold cross-validation คืออะไร? สมัครฟรี: https://sharpskill.dev/th/login ## หัวข้อสัมภาษณ์ Data Science & ML อื่นๆ - [พื้นฐาน Python](https://sharpskill.dev/th/technologies/data-science/interview-questions/python-basics.md): 25 คำถาม, Junior - [การเขียนโปรแกรมเชิงวัตถุด้วย Python](https://sharpskill.dev/th/technologies/data-science/interview-questions/python-oop.md): 20 คำถาม, Junior - [โครงสร้างข้อมูล Python](https://sharpskill.dev/th/technologies/data-science/interview-questions/python-data-structures.md): 20 คำถาม, Junior - [พื้นฐาน Git](https://sharpskill.dev/th/technologies/data-science/interview-questions/git-fundamentals.md): 18 คำถาม, Junior - [พื้นฐาน SQL](https://sharpskill.dev/th/technologies/data-science/interview-questions/sql-basics.md): 20 คำถาม, Junior - [พื้นฐาน NumPy](https://sharpskill.dev/th/technologies/data-science/interview-questions/numpy-fundamentals.md): 22 คำถาม, Junior - [พื้นฐาน Pandas](https://sharpskill.dev/th/technologies/data-science/interview-questions/pandas-basics.md): 22 คำถาม, Junior - [Jupyter & Google Colab](https://sharpskill.dev/th/technologies/data-science/interview-questions/jupyter-colab.md): 16 คำถาม, Junior - [SQL Joins และคิวรีขั้นสูง](https://sharpskill.dev/th/technologies/data-science/interview-questions/sql-joins-advanced.md): 22 คำถาม, Mid-Level - [Pandas ขั้นสูง](https://sharpskill.dev/th/technologies/data-science/interview-questions/pandas-advanced.md): 24 คำถาม, Mid-Level - [การแสดงผลข้อมูลด้วย Matplotlib & Seaborn](https://sharpskill.dev/th/technologies/data-science/interview-questions/matplotlib-seaborn.md): 20 คำถาม, Mid-Level - [การแสดงผลแบบโต้ตอบด้วย Plotly](https://sharpskill.dev/th/technologies/data-science/interview-questions/plotly-interactive.md): 18 คำถาม, Mid-Level - [สถิติเชิงพรรณนา](https://sharpskill.dev/th/technologies/data-science/interview-questions/statistics-descriptive.md): 20 คำถาม, Mid-Level - [สถิติเชิงอนุมาน](https://sharpskill.dev/th/technologies/data-science/interview-questions/statistics-inferential.md): 24 คำถาม, Mid-Level - [Web Scraping](https://sharpskill.dev/th/technologies/data-science/interview-questions/web-scraping.md): 18 คำถาม, Mid-Level - [BigQuery & Cloud Data](https://sharpskill.dev/th/technologies/data-science/interview-questions/bigquery-cloud.md): 18 คำถาม, Mid-Level - [Feature Engineering](https://sharpskill.dev/th/technologies/data-science/interview-questions/feature-engineering.md): 22 คำถาม, Mid-Level - [ML แบบมีผู้สอน: การถดถอย](https://sharpskill.dev/th/technologies/data-science/interview-questions/ml-supervised-regression.md): 24 คำถาม, Mid-Level - [ML แบบมีผู้สอน: การจำแนกประเภท](https://sharpskill.dev/th/technologies/data-science/interview-questions/ml-supervised-classification.md): 24 คำถาม, Mid-Level - [Decision Trees และ Ensembles](https://sharpskill.dev/th/technologies/data-science/interview-questions/ml-trees-ensembles.md): 24 คำถาม, Mid-Level - [Unsupervised ML](https://sharpskill.dev/th/technologies/data-science/interview-questions/ml-unsupervised.md): 22 คำถาม, Mid-Level - [Time Series และการพยากรณ์](https://sharpskill.dev/th/technologies/data-science/interview-questions/time-series-forecasting.md): 22 คำถาม, Mid-Level - [พื้นฐาน Deep Learning](https://sharpskill.dev/th/technologies/data-science/interview-questions/deep-learning-fundamentals.md): 24 คำถาม, Senior - [TensorFlow & Keras](https://sharpskill.dev/th/technologies/data-science/interview-questions/tensorflow-keras.md): 22 คำถาม, Senior - [CNN และการจำแนกภาพ](https://sharpskill.dev/th/technologies/data-science/interview-questions/cnn-image-classification.md): 24 คำถาม, Senior - [RNN และซีเควนซ์](https://sharpskill.dev/th/technologies/data-science/interview-questions/rnn-sequences.md): 22 คำถาม, Senior - [Transformers และ Attention](https://sharpskill.dev/th/technologies/data-science/interview-questions/transformers-attention.md): 24 คำถาม, Senior - [NLP และ Hugging Face](https://sharpskill.dev/th/technologies/data-science/interview-questions/nlp-huggingface.md): 24 คำถาม, Senior - [GenAI และ LangChain](https://sharpskill.dev/th/technologies/data-science/interview-questions/genai-langchain.md): 24 คำถาม, Senior - [MLOps และการ Deploy](https://sharpskill.dev/th/technologies/data-science/interview-questions/mlops-deployment.md): 24 คำถาม, Senior --- Source: SharpSkill (https://sharpskill.dev), tech interview preparation for your real stack. HTML version of this page: https://sharpskill.dev/th/technologies/data-science/interview-questions/ml-pipelines-validation