Mitigating Outlier Activations in Low-Precision Fine-Tuning of Language Models

Alireza Ghaffari; Justin Yu; Mahsa Nejad; Masoud Asgharian; Boxing Chen; Vahid Partovi Nia

Research.Publish.Connect.

*Please fill out at least one Field. *Value must be an number!

Title:
ISBN:
Year:
Acronym:
Subject:

Advanced Search Proceedings Search

If you're looking for an exact phrase use quotation marks on text fields.

*Please fill out at least one Field.

Title:
Author:
Affiliation:
Subject:

Advanced Search Papers Search

If you're looking for an exact phrase use quotation marks on text fields.

*Please fill out at least one Field.

Name:
Affiliation:
Country:
Conference:
Subject:

Advanced Search Authors Search

If you're looking for an exact phrase use quotation marks on text fields.

*Please fill out at least one Field.

Name:
Country:
Subject:

Advanced Search Affiliations Search

If you're looking for an exact phrase use quotation marks on text fields.

Proceedings

Proceedings Search *Please fill out at least one Field. *Value must be an number!

Title:
ISBN:
Year:
Acronym:
Subject:

Advanced Search Proceedings Search

If you're looking for an exact phrase use quotation marks on text fields.

Papers

Papers Search *Please fill out at least one Field.

Title:
Author:
Affiliation:
Subject:

Advanced Search Papers Search

If you're looking for an exact phrase use quotation marks on text fields.

Authors

Authors Search *Please fill out at least one Field.

Name:
Affiliation:
Country:
Conference:
Subject:

Advanced Search Authors Search

If you're looking for an exact phrase use quotation marks on text fields.

Advanced Search

Paper

Mitigating Outlier Activations in Low-Precision Fine-Tuning of Language Models

Topics: Deep Learning and Neural Networks; Natural Language Processing

In Proceedings of the 13th International Conference on Pattern Recognition Applications and Methods ICPRAM - Volume 1, 478-484, 2024 , Rome, Italy

Authors: Alireza Ghaffari ¹ ; Justin Yu ¹ ; Mahsa Nejad ¹ ; Masoud Asgharian ² ; Boxing Chen ¹ and Vahid Partovi Nia ¹

Affiliations: ¹ Huawei Noah’s Ark Lab, Montreal, Canada ; ² Department of Mathematics and Statistics, McGill University, Montreal, Canada

Keyword(s): Accelerated Training, Compressed Training, Low-Precision Fine-Tuning, Language Models.

Abstract: Low-precision fine-tuning of language models has gained prominence as a cost-effective and energy-efficient approach to deploying large-scale models in various applications. However, this approach is susceptible to the existence of outlier values in activation. The outlier values in the activation can negatively affect the performance of fine-tuning language models in the low-precision regime since they affect the scaling factor and thus make representing smaller values harder. This paper investigates techniques for mitigating outlier activation in low-precision integer fine-tuning of the language models. Our proposed novel approach enables us to represent the outlier activation values in 8-bit integers instead of floating-point ( FP16) values. The benefit of using integers for outlier values is that it enables us to use operator tiling to avoid performing 16-bit integer matrix multiplication to address this problem effectively. We provide theoretical analysis and supporting experime nts to demonstrate the effectiveness of our approach in improving the robustness and performance of low-precision fine-tuned language models. (More)

CC BY-NC-ND 4.0

Guest: Register as new SciTePress user now for free.

SciTePress user: please login.

My Papers

You are not signed in, therefore limits apply to your IP address 216.73.216.207

In the current month:

Recent papers: 100 available of 100 total

2⁺ years older papers: 200 available of 200 total

Paper citation in several formats:

Ghaffari, A., Yu, J., Nejad, M., Asgharian, M., Chen, B., Partovi Nia and V. (2024). Mitigating Outlier Activations in Low-Precision Fine-Tuning of Language Models. In Proceedings of the 13th International Conference on Pattern Recognition Applications and Methods - ICPRAM; ISBN 978-989-758-684-2; ISSN 2184-4313, SciTePress, pages 478-484. DOI: 10.5220/0012567700003654

@conference{icpram24,
author={Alireza Ghaffari and Justin Yu and Mahsa Nejad and Masoud Asgharian and Boxing Chen and Vahid {Partovi Nia}},
title={Mitigating Outlier Activations in Low-Precision Fine-Tuning of Language Models},
booktitle={Proceedings of the 13th International Conference on Pattern Recognition Applications and Methods - ICPRAM},
year={2024},
pages={478-484},
publisher={SciTePress},
organization={INSTICC},
doi={10.5220/0012567700003654},
isbn={978-989-758-684-2},
issn={2184-4313},
}

TY - CONF

JO - Proceedings of the 13th International Conference on Pattern Recognition Applications and Methods - ICPRAM
TI - Mitigating Outlier Activations in Low-Precision Fine-Tuning of Language Models
SN - 978-989-758-684-2
IS - 2184-4313
AU - Ghaffari, A.
AU - Yu, J.
AU - Nejad, M.
AU - Asgharian, M.
AU - Chen, B.
AU - Partovi Nia, V.
PY - 2024
SP - 478
EP - 484
DO - 10.5220/0012567700003654
PB - SciTePress