Publicación: Modelado y optimización del Trade-off Cómputo–Memoria mediante estrategias de rematerialización en entrenamiento de redes neuronales.
| dc.contributor.advisor | Torres Niño, Luis Alejandro | |
| dc.contributor.author | González Flores, Santiago | |
| dc.contributor.author | Leguizamó Céspedes, Andrey Felipe | |
| dc.contributor.evaluator | Vásquez Capacho, John William | |
| dc.contributor.evaluator | Benavides Arévalo, Bernardo Andrés | |
| dc.date.accessioned | 2026-08-21T16:22:21Z | |
| dc.date.created | 2026-08-14 | |
| dc.date.issued | 2026-08-14 | |
| dc.description.abstract | Durante el entrenamiento de redes neuronales profundas, la etapa de retropropagación exige retener en memoria las activaciones intermedias generadas en el forward pass, lo que provoca un consumo de memoria que restringe el tamaño de lote viable y, en dispositivos con recursos acotados, puede hacer inviable el entrenamiento de modelos de gran escala. La rematerialización de activaciones es una técnica que mitiga este problema descartando selectivamente dichas activaciones y recomputándolas cuando se necesitan durante la retropropagación, trasladando así la presión del sistema de memoria hacia el presupuesto de cómputo. Este trabajo desarrolló un estudio experimental y de modelado predictivo de ese compromiso. Se construyó un conjunto de datos de 2.064 configuraciones de entrenamiento, obtenidas mediante perfilado con PyTorch Profiler sobre seis arquitecturas de referencia (MLP-3, MLP-5, MLP-7, LeNet-5, AlexNet y VGG-16) ejecutadas en dos GPU de distinta familia (NVIDIA y AMD), empleando la infraestructura del supercomputador GUANE de la Unidad SC3–VIE de la UIS. Sobre estos datos se entrenaron el SGD Regressor y Random Forest que alcanzaron, en el conjunto de prueba, R2 = 0,8594 para el tiempo total y R2 = 0,8330 para el consumo de memoria. La validación mediante las pruebas estadísticas mostró que aplicar rematerialización produce diferencias significativas tanto en memoria como en tiempo frente al entrenamiento sin la técnica, mientras que la elección entre las variantes selectiva y uniforme no introduce diferencias estadísticamente distinguibles del ruido experimental. El análisis desagregado por arquitectura permitió establecer lineamientos concretos: en redes MLP y LeNet-5 la técnica no reporta beneficio; en AlexNet ofrece mejoras moderadas; y en VGG-16 reduce el consumo de memoria, con un sobrecoste computacional acotado. Estos resultados proveen criterios reproducibles para seleccionar y configurar estrategias de rematerialización en entornos académicos e industriales con restricciones de hardware. | |
| dc.description.abstractenglish | During the training of deep neural networks, the backpropagation stage requires storing the intermediate activations generated during the forward pass in memory, which results in memory consumption that limits the viable batch size and, on devices with limited resources, can make training large-scale models unfeasible. Activation rematerialization is a technique that mitigates this problem by selectively discarding these activations and recomputing them when needed during backpropagation, thereby shifting the pressure from the memory system to the computational budget. This study conducted an experimental and predictive modeling analysis of this trade-off. A dataset of 2,064 training configurations was constructed, obtained through profiling with the PyTorch Profiler on six reference architectures (MLP-3, MLP-5, MLP-7, LeNet-5, AlexNet, and VGG-16) running on two GPUs from different families (NVIDIA and AMD), using the infrastructure of the GUANE supercomputer at the SC3–VIE Unit of the UIS. The SGD Regressor and Random Forest were trained on this data and achieved, on the test set, R² = 0.8594 for total time and R² = 0.8330 for memory consumption. Validation through statistical tests showed that applying rematerialization produces significant differences in both memory and time compared to training without the technique, while the choice between the selective and uniform variants does not introduce differences statistically distinguishable from experimental noise. The analysis broken down by architecture allowed us to establish specific guidelines: in MLP and LeNet-5 networks, the technique offers no benefit; in AlexNet, it offers moderate improvements; and in VGG-16, it reduces memory consumption with a limited computational overhead. These results provide reproducible criteria for selecting and configuring rematerialization strategies in academic and industrial environments with hardware constraints. | |
| dc.description.degreelevel | Pregrado | |
| dc.description.degreename | Ingeniero de Sistemas | |
| dc.format.mimetype | application/pdf | |
| dc.identifier.instname | Universidad Industrial de Santander | |
| dc.identifier.reponame | Universidad Industrial de Santander | |
| dc.identifier.repourl | https://noesis.uis.edu.co | |
| dc.identifier.uri | https://noesis.uis.edu.co/handle/20.500.14071/48157 | |
| dc.language.iso | spa | |
| dc.publisher | Universidad Industrial de Santander | |
| dc.publisher.faculty | Facultad de Ingeníerias Fisicomecánicas | |
| dc.publisher.program | Ingeniería de Sistemas | |
| dc.publisher.school | Escuela de Ingeniería de Sistemas e Informática | |
| dc.rights | info:eu-repo/semantics/openAccess | |
| dc.rights.accessrights | info:eu-repo/semantics/openAccess | |
| dc.rights.coar | http://purl.org/coar/access_right/c_abf2 | |
| dc.rights.creativecommons | Atribución-NoComercial-SinDerivadas 4.0 Internacional (CC BY-NC-ND 4.0) | |
| dc.rights.license | Atribución-NoComercial-SinDerivadas 2.5 Colombia (CC BY-NC-ND 2.5 CO) | |
| dc.rights.uri | https://creativecommons.org/licenses/by-nc-nd/4.0/ | |
| dc.subject | redes neuronales profundas | |
| dc.subject | consumo de memoria | |
| dc.subject | rematerialización de activaciones | |
| dc.subject | gradient checkpointing | |
| dc.subject | optimización cómputo memoria | |
| dc.subject.keyword | deep neural networks | |
| dc.subject.keyword | memory consumption | |
| dc.subject.keyword | activation rematerialization | |
| dc.subject.keyword | gradient checkpointing | |
| dc.subject.keyword | compute-memory optimization | |
| dc.title | Modelado y optimización del Trade-off Cómputo–Memoria mediante estrategias de rematerialización en entrenamiento de redes neuronales. | |
| dc.title.english | Modeling and optimization of the Compute-Memory Trade-off through rematerialization strategies in neural network training | |
| dc.type.coar | http://purl.org/coar/resource_type/c_7a1f | |
| dc.type.hasversion | http://purl.org/coar/version/c_b1a7d7d4d402bcce | |
| dc.type.local | Tesis/Trabajo de grado - Monografía - Pregrado | |
| dspace.entity.type | Publication |
Archivos
Bloque original
1 - 4 de 4
Cargando...
- Nombre:
- Carta Autorización de Uso.pdf
- Tamaño:
- 156.4 KB
- Formato:
- Adobe Portable Document Format
Cargando...
- Nombre:
- Nota de proyecto.pdf
- Tamaño:
- 111.21 KB
- Formato:
- Adobe Portable Document Format
Bloque de licencias
1 - 1 de 1
Cargando...
- Nombre:
- license.txt
- Tamaño:
- 2.17 KB
- Formato:
- Item-specific license agreed to upon submission
- Descripción:
