New game
Download
Get Academic Plan
Share game
Integrate it into your platform

You can integrate the game into an LMS compatible with LTI 1.1 or LTI 1.3 such as Canvas, Moodle, or Blackboard. This way, the scores will be automatically saved into the platform’s gradebook.
Download
You have exceeded the maximum number of games you can integrate into Google Classroom with your current Plan.

To integrate as many games as you want in Google Classroom, you need an Academic Plan or a Commercial Plan.

You have exceeded the maximum number of games you can integrate into Microsoft Teams with your current Plan.

To integrate as many games as you want in Microsoft Teams, you need an Academic Plan or a Commercial Plan.

Downloading games is an exclusive feature for users with an Academic Plan or a Commercial Plan.

Get your Academic Plan or your Commercial Plan now and start integrating your games into your LMS, website or blog.

If you wish, you can download a demo game here and test its integration:

SPARK - LRBA 5

Quiz

(2)
Played 33 %Accuracy 86 Average time 07:16

About this activity

DataFrames y Datasets - Uso y Diferencias

Created by

Mexico

Download the paper version to play

Make your own free game from our game creator
Compete against your friends to see who gets the best score in this game

Top Games

%
Anonymous
Anonymous
%
%
%
You have exceeded the maximum number of games you can print with your current Plan.

To print as many games as you want, you need an Academic Plan or a Commercial Plan.

Print your game
SPARK - LRBA 5
 

SPARK - LRBA 5Online version

DataFrames y Datasets - Uso y Diferencias

by HR Mexico
1

¿Cuál es la diferencia principal entre DataFrame y Dataset en Scala?

2

¿Qué engine optimiza ambos?

3

En PySpark, ¿qué API usar para type-safety?

4

Overhead performance Datasets vs DataFrames?

5

Creación Dataset tipado requiere:

6

Mejor para MLlib pipelines:

7

AQE optimiza dinámicamente:

8

SaveMode por default:

9

Tungsten beneficio clave:

10

¿Cuándo usar Dataset sobre DataFrame?

11

Broadcast join condición:

12

Encoder rol en Dataset:

13

Caso fraude: mejor inicial:

14

PartitionBy aplica a:

15

Whole-stage codegen:

16

InferSchema=true overhead:

17

Java Dataset requiere:

18

Benchmark TPC-DS:

19

Buena práctica cache:

20

BucketBy beneficio:

Explicación

En Scala, DataFrame = Dataset[Row] untyped; Dataset[T] compile-time safe

Comparten Catalyst para plans, Tungsten para exec off-heap.

PySpark no soporta Dataset API typed

Datasets deserializan a JVM objects en map/filter.

df.as[Person] usa implicits encoder.

Spark MLlib usa DataFrame API.

Adaptive Query Execution post-shuffle

Lanza excepción si tabla existe.

Vectorized exec minimiza JVM overhead.

Compile-time safety nested cases.

Evita shuffle broadcast memoria nodos.

Permite Catalyst en typed ops.

Optimizado Catalyst para ETL joins.

Crea dirs por columna.

Tungsten genera código optimizado.

Explicación: Analiza sample datos; usa schema explícito.

POJO con getters para bean encoder.

Optimizaciones Catalyst/Tungsten.

MEMORY_AND_DISK por default spill.

Explicación: Hash-based distribución buckets.

Are you sure you want to leave the page?

If you leave the page, you will lose your game progress.