I don’t know whether this has been well enough praised here, but I wanted to once again (if it’s already been praised) heap praise upon Rupert Lane’s “Girdlock” tool (GitHub - rupertl/gridlock: Extract source code from line printer listing PDFs. · GitHub) which is designed specifically to to correct column-wise OCR from old printouts. The tool (operated, in these cases by Rupert himself) extracted the Logic Theorist from Stefferud’s 1963 RAND report, and recently extract the code for a 1620 IPL interpreter from about the same era from a PDF of the printout: iplv-listings/1620-oregon at main · rupertl/iplv-listings · GitHub