JUSTIN ZHANG - SWAN LAB
Happy to see the progression with IFS. I'll start recreating their scheduler this week.
Github Lab credentials are still needed to create the Github Organization.
Domain decision for the lab website is still pending.
Sent a request for a VM instance via the school service desk.
itsjira.carleton.ca/servicedeskSending out a message about a September lunch to coordinate a date with everyone.
Nothing research-relevant on the list today. Since your field is Software Engineering, I wanted to talk about versioning - specifically Calendar Versioning (CalVer).

CalVer tags releases by date instead of arbitrary semantic numbers.
Two foundational papers on training large models across multiple accelerators: GPipe (pipeline parallelism) and Megatron-LM (tensor parallelism).
Enables training on multiple accelerators (GPUs, TPUs) in a parallel manner. What makes this paper unique:

A batch-splitting algorithm overlaps compute by splitting the forward and backward pass into smaller micro-batches.
On backpropagation, each weight w is provided a single gradient telling it which direction to push.
Each weight accumulates a collection of gradients across micro-batches:
w_1 + w_2 + w_3 + ... + w_nDrop-in library partitions a model > splits batches into micro-batches > pipelines them across accelerators > accumulates gradients > syncs.
Not relevant - hardware scaling for multi-GPU, and only applicable for training, not inference. Though many papers build on micro-batching for inference.
Just like last time when we discussed FlashAttention and how matrices are sliced up, Megatron also uses partitioned matrices (block matrix wiki page).
FlashAttention slices a matrix into sub-matrices to fit into L1 caches (SRAM) for IO-aware computation. Megatron slices the matrix so the multiplications can run on different accelerators, then concatenates them.

Tiling, block matrix multiplication, and partitioned matrix multiplication all mean the same thing here.
Not relevant - also hardware scaling for multi-GPU. Where GPipe revolves around how batches move through the network, Megatron changes the shape of each batch by slicing the matrix multiplications themselves across accelerators.
Both contribute to training at scale - GPipe via how batches flow, Megatron via how matrices are shaped. Neither touches inference scheduling, my research direction.

The orchestrator scores tasks and hands work to workers as they poll.
Right now the orchestrator has “amnesia” - every time a worker polls, it computes the score from scratch and decides which task to hand over.
In the current score formulation, 0.1 and 0.9 are hardcoded weights:
score = 0.1 * (Norm. backlog) + 0.9 * (Norm. wait time)DES already logs every scheduling decision it makes, so we could train an ML model on those historical logs without ever touching production. (This depends on whether they are open to sharing this data - if not, I can synthesize it, but it won't be realistic.)
Can an RL agent look at these logs and find better weights than 0.1 and 0.9?
You mentioned ML task prediction last meeting. Currently “wait time” just means how long a task has been sitting in the queue - but if we could predict how long a task will actually take, we could schedule smarter (e.g. priority to short-but-waited jobs).
Note: a small model could add overhead since the score is recomputed on each worker poll - it needs to be really, really fast.JUSTIN ZHANG - SWAN LAB - MCS MEETING AUG 13