# PySpark - 大規模処理 (Data Engineering) > SparkSession、RDD vs DataFrame、transformations、actions、partitioning、broadcast variables、UDFs、Spark SQL、caching - 20 面接問題 - Senior - [面接問題: Data Engineering](https://sharpskill.dev/ja/technologies/data-engineering/interview-questions.md) ## 1. PySparkアプリケーションを作成するための主要なエントリーポイントは何ですか? **回答** SparkSessionはSpark 2.0で導入された統一されたエントリーポイントです。古いSparkContext、SQLContext、HiveContextを単一のオブジェクトに置き換えます。SparkSessionを使用すると、DataFrameの作成、SQLクエリの実行、Sparkアプリケーションの集中管理が可能になります。 ## 2. PySparkにおけるRDDとDataFrameの基本的な違いは何ですか? **回答** DataFrameは名前付きで型付けされた列を持つ構造化スキーマを持ち、SparkがCatalystを介してクエリを最適化できます。RDDは構造化されていない分散コレクションで、Sparkは内部のデータ構造を知らないため、可能な最適化が制限されます。 ## 3. PySparkにおけるtransformationとactionの違いは何ですか? **回答** transformationは遅延評価され、計算をトリガーすることなく実行プランを構築します。actionはクラスター上でプランの実際の実行をトリガーし、結果をdriverに返します。この区別により、Sparkは実行前にプランを最適化できます。 ## さらに17問利用可能 - 次の操作のうち、PySparkのactionはどれですか? - PySparkでParquetファイルからDataFrameを作成するにはどうすればよいですか? 無料で登録: https://sharpskill.dev/ja/login ## その他のData Engineering面接トピック - [Linux & Shell - 基礎](https://sharpskill.dev/ja/technologies/data-engineering/interview-questions/linux-shell-basics.md): 20問, Junior - [Git & GitHub - 基礎](https://sharpskill.dev/ja/technologies/data-engineering/interview-questions/git-github-fundamentals.md): 20問, Junior - [データエンジニアリングのための高度なPython](https://sharpskill.dev/ja/technologies/data-engineering/interview-questions/python-advanced-de.md): 25問, Junior - [Docker - 基礎](https://sharpskill.dev/ja/technologies/data-engineering/interview-questions/docker-fundamentals.md): 25問, Junior - [Google Cloud Platform - 基礎](https://sharpskill.dev/ja/technologies/data-engineering/interview-questions/gcp-fundamentals.md): 20問, Junior - [CI/CDとコード品質](https://sharpskill.dev/ja/technologies/data-engineering/interview-questions/ci-cd-code-quality.md): 20問, Mid-Level - [Docker Compose](https://sharpskill.dev/ja/technologies/data-engineering/interview-questions/docker-compose.md): 20問, Mid-Level - [FastAPI - データAPI](https://sharpskill.dev/ja/technologies/data-engineering/interview-questions/fastapi.md): 20問, Mid-Level - [Data Engineering向けの高度なSQL](https://sharpskill.dev/ja/technologies/data-engineering/interview-questions/sql-advanced-de.md): 20問, Mid-Level - [Data Lake - アーキテクチャと取り込み](https://sharpskill.dev/ja/technologies/data-engineering/interview-questions/data-lake.md): 20問, Mid-Level - [データエンジニアリングのためのBigQuery](https://sharpskill.dev/ja/technologies/data-engineering/interview-questions/bigquery-de.md): 20問, Mid-Level - [PostgreSQL - 管理](https://sharpskill.dev/ja/technologies/data-engineering/interview-questions/postgresql-admin.md): 20問, Mid-Level - [Data EngineeringのためのData Modeling](https://sharpskill.dev/ja/technologies/data-engineering/interview-questions/data-modeling-de.md): 20問, Mid-Level - [Fivetran & Airbyte - データ取り込み](https://sharpskill.dev/ja/technologies/data-engineering/interview-questions/fivetran-airbyte.md): 20問, Mid-Level - [dbt - 基礎](https://sharpskill.dev/ja/technologies/data-engineering/interview-questions/dbt-fundamentals-de.md): 20問, Mid-Level - [Apache Airflow - 基礎](https://sharpskill.dev/ja/technologies/data-engineering/interview-questions/airflow-fundamentals.md): 20問, Mid-Level - [Kubernetes - 基礎](https://sharpskill.dev/ja/technologies/data-engineering/interview-questions/kubernetes-fundamentals.md): 20問, Mid-Level - [dbt - 高度な機能](https://sharpskill.dev/ja/technologies/data-engineering/interview-questions/dbt-advanced-de.md): 20問, Senior - [ETL / ELT / ETLT パターン](https://sharpskill.dev/ja/technologies/data-engineering/interview-questions/etl-elt-patterns.md): 20問, Senior - [Apache Airflow - 上級](https://sharpskill.dev/ja/technologies/data-engineering/interview-questions/airflow-advanced.md): 20問, Senior - [Airflow + dbt - パイプラインオーケストレーション](https://sharpskill.dev/ja/technologies/data-engineering/interview-questions/airflow-dbt-integration.md): 20問, Senior - [Google Pub/Sub - データストリーミング](https://sharpskill.dev/ja/technologies/data-engineering/interview-questions/pubsub-streaming.md): 20問, Senior - [Apache Beam & Dataflow](https://sharpskill.dev/ja/technologies/data-engineering/interview-questions/apache-beam-dataflow.md): 20問, Senior - [Kubernetes - 本番環境とスケーリング](https://sharpskill.dev/ja/technologies/data-engineering/interview-questions/kubernetes-advanced.md): 20問, Senior - [Terraform - Infrastructure as Code](https://sharpskill.dev/ja/technologies/data-engineering/interview-questions/terraform.md): 20問, Senior - [NoSQLデータベース](https://sharpskill.dev/ja/technologies/data-engineering/interview-questions/nosql-databases.md): 20問, Senior - [モダンなData Architecture](https://sharpskill.dev/ja/technologies/data-engineering/interview-questions/data-architecture.md): 20問, Senior - [モニタリングとオブザーバビリティ](https://sharpskill.dev/ja/technologies/data-engineering/interview-questions/monitoring-observability.md): 20問, Senior - [IAMとデータセキュリティ](https://sharpskill.dev/ja/technologies/data-engineering/interview-questions/iam-security-de.md): 20問, Senior --- Source: SharpSkill (https://sharpskill.dev), tech interview preparation for your real stack. HTML version of this page: https://sharpskill.dev/ja/technologies/data-engineering/interview-questions/pyspark