Building Flash Attention from Source

Notes from a few rounds of compiling Flash Attention on an A800 server. I walk through checking that ninja is actually working, limiting the build to the GPU architecture I need, and balancing parallel jobs against RAM usage. If your build takes forever or ends with a “killed” error, these are the settings and pitfalls I wish I had checked first.

      

Install CUDA and an NLP Stack with Conda (No Root)

No root access on the remote server, but still want a newer PyTorch and transformers setup? These are my notes on using Conda to install CUDA and the tools I need without changing the system driver. I cover version choices, checking that PyTorch can use the GPU, pointing programs to the Conda CUDA installation, and watching out for xformers dependency changes.