How many FLOPS will be used to train GPT-4 (if it is released)?
Resolved
2.1×10²⁵
Forecast Timeline
- 1d 1w 2m all
- Aug 24
- Aug 26
- Aug 28
- Aug 30
- Sep 01
- Sep 03
- Sep 05
- Sep 07
- Sep 09
- Sep 11
- Sep 13
- Sep 15
- Sep 17
- Sep 19
- Sep 21
- Sep 23
- Sep 25
- Sep 27
- Sep 29
- Oct 01
- Oct 03
- Oct 05
- Oct 07
- Oct 09
- Oct 11
- Oct 13
- Oct 15
- Oct 17
- Oct 19
- Oct 21
- Oct 23
Key Factors (0)
No key factors yet Add some that might influence this forecast.
Comments
NMorrison
- Resolved as 2.1e+25 based on Epoch's estimate, which seems to be the best data available. I’ve resolved this as of Oct 23, 2023, which is the date Epoch published their Parameter, Compute and Data Trends database.
qumeric
- Prediction: 1×10²⁵ (4.5×10²⁴ - 2.5×10²⁵)
- Epoch estimated 2.2e25 (their 90% CI is 1e25-5.2e25).
citizen
- Should the method for predicting this question about a LLM from OpenAI be different from the method for predicting the other GPT-4 question, which specified a year? If so, how should the methodology differ?
MayMeta
- From OpenAI's GPT-4 report: "Given both the competitive landscape and the safety implications of large-scale models like GPT-4, this report contains no further details about the architecture (including model size), hardware, training compute, dataset construction, training method, or similar."
ingrahammark7
- Prediction: 1.6×10²⁴ (9.1×10²³ - 2.9×10²⁴) - E23 trained gpt3. The largest supercomputer could do e24 per year.
Trunton
- According to a witness who allegedly has a friend with access to GPT-4, "it is just as exciting a Leap as GPT-3 was."
Anthony
- All I can say is thank you for not using petaflop/s*days as your unit.
Tamay
- GPT-1 used 1.10E19, GPT-2 used 2.49E21, and GPT-3 used 3.14E23 FLOPS. Seem like a stable progression of 2OOMs/GPT. I expect that GPT-4 will come in at 8E23 to 2E25 FLOPS given the current global chip shortage.
Tamay
- Some reference points:
- GPT-3 took 3.14E+23 FLOPS to train
- Deepmind's GOPHER took 6.31E+23 FLOPS to train
- The largest disclosed ML experiment to date (Megatron-Turing NLG 530B) took 1.35E+24 to train.
- Some reference points: