MemFLoRA: Memory-Floor LoRA for CNN Adaptation at the Edge
Mehmet Emre Akbulut, Johannes Geier, Ulf Schlichtmann
- Published
- Oct 6, 2026 — 16:51 UTC
Problem
The paper addresses the challenge of activation state retention during the backward pass in Convolutional Neural Network (CNN) adaptation, particularly in resource-constrained environments such as edge devices. The authors propose a solution to this problem through a novel architecture, MemFLoRA, which is designed to optimize memory usage without compromising performance. This work is presented as a preprint and has not yet undergone peer review.
Method
The core technical contribution is the Memory-Floor LoRA (MemFLoRA), which employs a memory-first design principle. The architecture is characterized by the following components:
- Activation-Memory-Floor Criterion: This criterion allows for trainable backward computations that do not rely on full-width layer inputs, thereby reducing memory overhead.
- Freezing Down-Projection: The model freezes the down-projection component, which helps in minimizing memory usage during training.
- Training Scale-Matched Up-Projection: The up-projection component is trained to match the scale of the down-projection, ensuring efficient adaptation.
- Combining Eval-Mode Backbone Normalization: The approach integrates evaluation-mode backbone normalization with activation-minimal backward rules to further optimize memory usage.
The model is evaluated on three Human Activity Recognition (HAR) datasets, utilizing two unspecified CNN backbones for the experiments.
Results
The results demonstrate significant memory savings compared to full fine-tuning approaches:
- Saved-Activation Memory Reduction: 98.5-98.7% reduction in memory usage.
- Peak Training-State Memory Reduction: 94.9-97.3% reduction in peak training-state memory.
- Performance: MemFLoRA matches or exceeds performance metrics of existing CNN Parameter-Efficient Fine-Tuning (PEFT) baselines, although specific performance metrics are not disclosed in the text.
Limitations
The authors do not report any limitations in the study. However, the lack of detailed performance metrics and the unspecified nature of the CNN backbones may limit the generalizability of the findings.
Why it matters
The implications of this work are significant for the deployment of CNNs in edge computing scenarios, where memory and computational resources are often limited. By drastically reducing memory requirements while maintaining or improving performance, MemFLoRA enables more efficient model adaptation, potentially broadening the applicability of CNNs in real-time applications such as human activity recognition.
By Turing Wire Research Desk · Oct 6, 2026 · How we work →
Summarised from the paper by the Turing Wire Research Desk. The full paper has the complete methods and results.
Source: arXiv cs.AI
