dask chunking tutorial outline #157

New issue

Open

Labels

enhancement

@dcherian

Description

@dcherian

dcherian

opened

on Jan 31, 2023

from the pangeo working meeting discussion with @mgrover1 @jmunroe @norlandrhagen

Here's an outline for an intermediate tutorial talking about dask chunking specifically for Xarray users

Motivation: why care about chunk size?

demonstrate relation between chunk size and computation time / number of tasks with a simple example?
- maybe even memory usage
https://tutorial.dask.org/02_array.html#Choosing-good-chunk-sizes
https://docs.dask.org/en/stable/array-chunks.html

Keeping track

monitoring chunk sizes and num tasks throughout the pipeline using the repr
- use some images
while output blocks may be small (say after a big reduction), intermediate blocks need not be.
So keep monitoring chunksizes (and tasks) throughout the pipeline.

Why is it important to choose appropriate chunks early in the pipeline?

Demonstrate that rechunking is not cheap in most cases

Specify chunks when reading data

Avoid chunks="auto".
- https://docs.dask.org/en/stable/array-chunks.html#automatic-chunking
Specifying chunks during data read
- open_dataset
- open_mfdataset
Analysis vs storage chunks:
- Dask chunks should be a multiple of chunks on disk
- talk about aligning chunks with files stored on disk
- @djhoese example

Metadata

Assignees

No one assigned

Labels

enhancement

Type

No type

Projects

No projects

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Navigation Menu

Search code, repositories, users, issues, pull requests...

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Uh oh!

dask chunking tutorial outline #157

Description

Motivation: why care about chunk size?

Keeping track

Why is it important to choose appropriate chunks early in the pipeline?

Specify chunks when reading data

Metadata

Metadata

Assignees

Labels

Type

Projects

Milestone

Relationships

Development

Issue actions