New game
Download
Get Academic Plan
Share game
Integrate it into your platform

You can integrate the game into an LMS compatible with LTI 1.1 or LTI 1.3 such as Canvas, Moodle, or Blackboard. This way, the scores will be automatically saved into the platform’s gradebook.
Download
You have exceeded the maximum number of games you can integrate into Google Classroom with your current Plan.

To integrate as many games as you want in Google Classroom, you need an Academic Plan or a Commercial Plan.

You have exceeded the maximum number of games you can integrate into Microsoft Teams with your current Plan.

To integrate as many games as you want in Microsoft Teams, you need an Academic Plan or a Commercial Plan.

Downloading games is an exclusive feature for users with an Academic Plan or a Commercial Plan.

Get your Academic Plan or your Commercial Plan now and start integrating your games into your LMS, website or blog.

If you wish, you can download a demo game here and test its integration:

SPARK - LRBA 1

Quiz

(4)
Played 44 %Accuracy 70 Average time 07:22

About this activity

Introducción a Apache Spark

Created by

Mexico

Download the paper version to play

Make your own free game from our game creator
Compete against your friends to see who gets the best score in this game

Top Games

%
Anonymous
Anonymous
%
%
%
You have exceeded the maximum number of games you can print with your current Plan.

To print as many games as you want, you need an Academic Plan or a Commercial Plan.

Print your game
SPARK - LRBA 1
 

SPARK - LRBA 1Online version

Introducción a Apache Spark

by HR Mexico
1

¿Cuál es la diferencia principal entre transformaciones y acciones en RDD?

2

En Spark, ¿qué optimizador maneja DataFrames/SQL?

3

¿Qué StorageLevel usa memoria y disco serializado?

4

En Structured Streaming, ¿qué maneja late data?

5

¿Cuál evita shuffle en groupByKey?

6

¿Feature de Spark 3.0+ para skew?

7

¿API unificada post-Spark 2.0?

8

En arquitectura, ¿coordina tasks?

9

¿Para joins pequeños en DataFrames?

10

¿Fault-tolerance vía?

11

¿Modo output para agregaciones streaming completas?

12

¿Tamaño ideal partición?

13

¿Wide dependency genera?

14

En Netflix, Spark reemplaza?

15

¿Off-heap memory?

16

¿Para ML pipelines?

17

¿Dynamic allocation requiere?

18

En Uber, Spark procesa?

19

¿Whole-stage codegen?

20

¿Cluster manager en cloud?

Explicación

Transformaciones construyen DAG lazy; acciones materializan cómputo.

Catalyst aplica rule/cost-based optimizaciones pre-ejecución.

MEMORY_AND_DISK_SER serializa para ahorrar memoria.

Watermark define threshold para dropear eventos tardíos.

reduceByKey combina localmente antes de shuffle.

Adaptive Query Execution ajusta runtime (skew handling).

Unifica batch/streaming/ML con optimizaciones.

Task Scheduler asigna tasks considerando locality.

Broadcast envía tabla a todos executors, evita shuffle.

Lineage permite recomputar particiones perdidas.

complete emite tabla completa por trigger.

Balancea overhead/paralelismo (Spark doc).

Wide requiere shuffle, divide stages.

10x speedup en ETL batch.

Tungsten usa off-heap para GC efficiency.

MLlib provee transformers/estimators distribuidos.

Escala executors dinámicamente.

Structured Streaming para movilidad real-time.

Fusiona stages en single function para velocidad.

Soporta YARN, Kubernetes para orquestación.

Are you sure you want to leave the page?

If you leave the page, you will lose your game progress.