arxiv:2503.14125

Frac-Connections: Fractional Extension of Hyper-Connections

Published on Mar 18

· Submitted by

mathfinder on Mar 19

Upvote

Authors:

Defa Zhu ,

Hongzhi Huang ,

Jundong Zhou ,

Zihao Huang ,

Yutao Zeng ,

Banggu Wu ,

Qiyang Min ,

Abstract

Residual connections are central to modern deep learning architectures, enabling the training of very deep networks by mitigating gradient vanishing. Hyper-Connections recently generalized residual connections by introducing multiple connection strengths at different depths, thereby addressing the seesaw effect between gradient vanishing and representation collapse. However, Hyper-Connections increase memory access costs by expanding the width of hidden states. In this paper, we propose Frac-Connections, a novel approach that divides hidden states into multiple parts rather than expanding their width. Frac-Connections retain partial benefits of Hyper-Connections while reducing memory consumption. To validate their effectiveness, we conduct large-scale experiments on language tasks, with the largest being a 7B MoE model trained on up to 3T tokens, demonstrating that Frac-Connections significantly outperform residual connections.

View arXiv page View PDF Add to collection

Community

mathfinder

Paper author Paper submitter 1 day ago

pengxiang

1 day ago

Congratulations, I really like this series of work.

I would like to ask if the authors have plans to open-source the weights :)

mathfinder

Paper author Paper submitter 1 day ago

Thank you very much for your kind words! We’re glad you like our work. Yes, we do have plans to open-source the weights, and we will announce it once everything is ready. Stay tuned!