JUSTIN ZHANG - CARLETON UNIVERSITY, SWAN LAB
Infrastructure is on hold for now. My effort is focused on Point 1 of your email: designing and building a per-task-type cost model.
Points 2, 3, and 4 are placed on hold. Current efforts are on Points 1, 5, and 6.
Messy natural language in (any feature space) → typesafe, structured JSON out - sub-second, and it cannot hallucinate. Trained with RLCD, introduced by Typesafe AI.
Every interface returns a probability or a distribution over an answer space you supply.
Pick one option from a set of answers you supply it.
NO ORDERA single probability answering a yes/no question.
PROBABILITYRate something against ordinal levels you supply it.
HAS ORDERThe cost grid uses Score: six ordinal levels from 0.01 GB-s to 1000 GB-s for every file/task combination.
SIZES: 1KB · 100KB · 1MB · 10MB · 100MB · 1GB · 10GB · 1TBThe purpose of the experiment: check whether these models provide a meaningful heuristic as a per-task-type cost model.
Does it improve fairness? We don't know yet - but we will get there :)
The full 5,896-row grid lives in an interactive HTML viewer. Open it preloaded with a model - then click any cell to see its full probability distribution.
Open-Jev-2B is the only model that orders expected cost by file size - System One Models feel like a favorable direction.
| MODEL | 1KB | 100KB | 1MB | 10MB | 100MB | 1GB | 10GB | 1TB |
|---|---|---|---|---|---|---|---|---|
| OPEN-JEV-2B | 1.31 | 1.76 | 2.29 | 2.70 | 2.84 | 3.28 | 3.11 | 3.12 |
| LAYA-421M | 2.15 | 2.32 | 2.25 | 2.34 | 2.37 | 2.38 | 2.39 | 2.54 |
| OPEN-JEV-DEBERTA | 2.34 | 2.33 | 2.34 | 2.34 | 2.38 | 2.45 | 2.47 | 2.59 |
MEAN EXPECTED LEVEL PER FILE SIZE · 6 ORDINAL BUCKETS · 0 = 0.01 GB-S · 5 = 1000 GB-S · 5,896 ROWS PER MODELPromising - but I will only be confident after a true benchmark and real variation along the task axis.
NONE OF THOSE LOOK LIKE ASYNCHRONOUS SCHEDULINGMEAN E IS FLAT ACROSS ALL 8 FILE SIZES (2.34 - 2.59)FILE · METADATA · GB-S · DURATION · MEMORY USAGE · CPU/GPU USAGEWATCH FOR MODEL DRIFT - FILETYPES CHANGEI WOULD LOVE TO INJECT FILE ENTROPY AS A FEATUREIF SUMMER-SCHEDULE MEETINGS CONTINUE, I WOULD LOVE TO DISCUSS MORE LITERATURESTILL ADJUSTING THE LIT REVIEW LISTAny suggestions on which next steps to prioritize?
Feedback given by IFS on this work - the answer to the ask.
EXPENSIVE INITIALLY - SAVES FUTURE COSTS (MEMOIZATION)INTUITIVELY, IT SEEMS LIKE LATENCY WILL INCREASEALSO PREFERS WE DON'T GO IN THE DIRECTION OF STEP 3 (EVALUATE NLI CLASSIFIERS)JUSTIN ZHANG - CARLETON UNIVERSITY, SWAN LAB