# PySpark - Large-Scale Processing (Data Engineering) > SparkSession, RDD vs DataFrame, transformations, actions, partitioning, broadcast variables, UDFs, Spark SQL, caching - 20 interview questions - Senior - [Interview Questions: Data Engineering](https://sharpskill.dev/en/technologies/data-engineering/interview-questions.md) ## 1. What is the main entry point for creating a PySpark application? **Answer** SparkSession is the unified entry point introduced in Spark 2.0. It replaces the old SparkContext, SQLContext, and HiveContext with a single object. SparkSession allows creating DataFrames, executing SQL queries, and configuring the Spark application in a centralized way. ## 2. What is the fundamental difference between an RDD and a DataFrame in PySpark? **Answer** A DataFrame has a structured schema with named and typed columns, allowing Spark to optimize queries through Catalyst. An RDD is an unstructured distributed collection where Spark doesn't know the internal data structure, limiting possible optimizations. ## 3. What is the difference between a transformation and an action in PySpark? **Answer** Transformations are lazily evaluated and build an execution plan without triggering computation. Actions trigger the actual execution of the plan on the cluster and return a result to the driver. This distinction allows Spark to optimize the plan before execution. ## 17 more questions available - Among the following operations, which one is a PySpark action? - How to create a DataFrame from a Parquet file in PySpark? Sign up for free: https://sharpskill.dev/en/login ## Other Data Engineering interview topics - [Linux & Shell - Fundamentals](https://sharpskill.dev/en/technologies/data-engineering/interview-questions/linux-shell-basics.md): 20 questions, Junior - [Git & GitHub - Fundamentals](https://sharpskill.dev/en/technologies/data-engineering/interview-questions/git-github-fundamentals.md): 20 questions, Junior - [Advanced Python for Data Engineering](https://sharpskill.dev/en/technologies/data-engineering/interview-questions/python-advanced-de.md): 25 questions, Junior - [Docker - Fundamentals](https://sharpskill.dev/en/technologies/data-engineering/interview-questions/docker-fundamentals.md): 25 questions, Junior - [Google Cloud Platform - Fundamentals](https://sharpskill.dev/en/technologies/data-engineering/interview-questions/gcp-fundamentals.md): 20 questions, Junior - [CI/CD and Code Quality](https://sharpskill.dev/en/technologies/data-engineering/interview-questions/ci-cd-code-quality.md): 20 questions, Mid-Level - [Docker Compose](https://sharpskill.dev/en/technologies/data-engineering/interview-questions/docker-compose.md): 20 questions, Mid-Level - [FastAPI - Data APIs](https://sharpskill.dev/en/technologies/data-engineering/interview-questions/fastapi.md): 20 questions, Mid-Level - [Advanced SQL for Data Engineering](https://sharpskill.dev/en/technologies/data-engineering/interview-questions/sql-advanced-de.md): 20 questions, Mid-Level - [Data Lake - Architecture and Ingestion](https://sharpskill.dev/en/technologies/data-engineering/interview-questions/data-lake.md): 20 questions, Mid-Level - [BigQuery for Data Engineering](https://sharpskill.dev/en/technologies/data-engineering/interview-questions/bigquery-de.md): 20 questions, Mid-Level - [PostgreSQL - Administration](https://sharpskill.dev/en/technologies/data-engineering/interview-questions/postgresql-admin.md): 20 questions, Mid-Level - [Data Modeling for Data Engineering](https://sharpskill.dev/en/technologies/data-engineering/interview-questions/data-modeling-de.md): 20 questions, Mid-Level - [Fivetran & Airbyte - Data Ingestion](https://sharpskill.dev/en/technologies/data-engineering/interview-questions/fivetran-airbyte.md): 20 questions, Mid-Level - [dbt - Fundamentals](https://sharpskill.dev/en/technologies/data-engineering/interview-questions/dbt-fundamentals-de.md): 20 questions, Mid-Level - [Apache Airflow - Fundamentals](https://sharpskill.dev/en/technologies/data-engineering/interview-questions/airflow-fundamentals.md): 20 questions, Mid-Level - [Kubernetes - Fundamentals](https://sharpskill.dev/en/technologies/data-engineering/interview-questions/kubernetes-fundamentals.md): 20 questions, Mid-Level - [dbt - Advanced Features](https://sharpskill.dev/en/technologies/data-engineering/interview-questions/dbt-advanced-de.md): 20 questions, Senior - [ETL / ELT / ETLT Patterns](https://sharpskill.dev/en/technologies/data-engineering/interview-questions/etl-elt-patterns.md): 20 questions, Senior - [Apache Airflow - Advanced](https://sharpskill.dev/en/technologies/data-engineering/interview-questions/airflow-advanced.md): 20 questions, Senior - [Airflow + dbt - Pipeline Orchestration](https://sharpskill.dev/en/technologies/data-engineering/interview-questions/airflow-dbt-integration.md): 20 questions, Senior - [Google Pub/Sub - Data Streaming](https://sharpskill.dev/en/technologies/data-engineering/interview-questions/pubsub-streaming.md): 20 questions, Senior - [Apache Beam & Dataflow](https://sharpskill.dev/en/technologies/data-engineering/interview-questions/apache-beam-dataflow.md): 20 questions, Senior - [Kubernetes - Production and Scaling](https://sharpskill.dev/en/technologies/data-engineering/interview-questions/kubernetes-advanced.md): 20 questions, Senior - [Terraform - Infrastructure as Code](https://sharpskill.dev/en/technologies/data-engineering/interview-questions/terraform.md): 20 questions, Senior - [NoSQL Databases](https://sharpskill.dev/en/technologies/data-engineering/interview-questions/nosql-databases.md): 20 questions, Senior - [Modern Data Architecture](https://sharpskill.dev/en/technologies/data-engineering/interview-questions/data-architecture.md): 20 questions, Senior - [Monitoring and Observability](https://sharpskill.dev/en/technologies/data-engineering/interview-questions/monitoring-observability.md): 20 questions, Senior - [IAM and Data Security](https://sharpskill.dev/en/technologies/data-engineering/interview-questions/iam-security-de.md): 20 questions, Senior --- Source: SharpSkill (https://sharpskill.dev), tech interview preparation for your real stack. HTML version of this page: https://sharpskill.dev/en/technologies/data-engineering/interview-questions/pyspark