New game
Download
Get Academic Plan
Share game
Integrate it into your platform

You can integrate the game into an LMS compatible with LTI 1.1 or LTI 1.3 such as Canvas, Moodle, or Blackboard. This way, the scores will be automatically saved into the platform’s gradebook.
Download
You have exceeded the maximum number of games you can integrate into Google Classroom with your current Plan.

To integrate as many games as you want in Google Classroom, you need an Academic Plan or a Commercial Plan.

You have exceeded the maximum number of games you can integrate into Microsoft Teams with your current Plan.

To integrate as many games as you want in Microsoft Teams, you need an Academic Plan or a Commercial Plan.

Downloading games is an exclusive feature for users with an Academic Plan or a Commercial Plan.

Get your Academic Plan or your Commercial Plan now and start integrating your games into your LMS, website or blog.

If you wish, you can download a demo game here and test its integration:

SPARK - LRBA 3

Quiz

(3)
Played 34 %Accuracy 80 Average time 08:09

About this activity

Spark SQL - Manejo y Consulta de Datos Estructurados

Created by

Mexico

Download the paper version to play

Make your own free game from our game creator
Compete against your friends to see who gets the best score in this game

Top Games

%
Anonymous
Anonymous
%
%
%
You have exceeded the maximum number of games you can print with your current Plan.

To print as many games as you want, you need an Academic Plan or a Commercial Plan.

Print your game
SPARK - LRBA 3
 

SPARK - LRBA 3Online version

Spark SQL - Manejo y Consulta de Datos Estructurados

by HR Mexico
1

¿Cuál es el rol principal de Catalyst Optimizer en Spark SQL?

2

En lectura de Parquet, ¿qué opción habilita predicate pushdown?

3

¿Cuál es la sintaxis para particionar al escribir un DataFrame?

4

¿Qué tipo de datos representa arrays anidados en esquemas?

5

En joins, ¿cuándo usar broadcast hint?

6

¿Qué función window computa ranking?

7

¿Qué habilita ACID en Spark SQL?

8

En streaming, ¿qué maneja datos tardíos?

9

¿Configuración activa AQE?

10

¿Buena práctica para UDFs de alto rendimiento?

11

¿Qué comando muestra plan de ejecución detallado?

12

En bucketing, ¿beneficia principalmente?

13

¿Ventaja de Datasets sobre DataFrames?

14

¿Qué evita múltiples filtros en cadena?

15

En Hive tables, ¿qué crea particiones dinámicas?

16

¿Métrica clave en Spark UI para skew?

17

¿Función para flatten arrays?

18

¿Configuración para partitions output?

19

En retail analytics, ¿particionamiento óptimo?

20

¿Extensión para metadatos persistentes?

Explicación

Catalyst aplica reglas como constant folding y join reordering para optimizar planes de ejecución.

Parquet soporta filtros en metadata, empujados antes de lectura completa

partitionBy crea directorios jerárquicos por valores de columna.

ArrayType(StructType()) para listas de objetos complejos.

Evita shuffle enviando tabla a todos ejecutores.

rank() asigna rangos dentro de particiones ordenadas.

Delta proporciona transacciones atómicas sobre Parquet.

withWatermark descarta datos fuera de ventana temporal.

AQE optimiza dinámicamente basado en estadísticas runtime.

Pandas UDFs aplican funciones a batches Pandas, acelerando 100x.

explain(extended) revela DAG con optimizaciones.

BucketBy hash-particiona para joins directos.

Datasets ofrecen chequeo estático de tipos.

Optimizador fusiona condiciones en where().

PARTITIONED BY habilita inserts dinámicos.

Desbalance en shuffle indica skew en keys.

explode() expande arrays a rows.

Controla tamaño y coalescing dinámico.

Prune por tiempo/localidad acelera queries comunes.

Hive Metastore persiste esquemas cross-sesiones.

Are you sure you want to leave the page?

If you leave the page, you will lose your game progress.