Module 13: CICS Recovery, Restart and Production Support
Emergency Restart
An emergency restart runs automatically when CICS is restarted after an abnormal end - a region abend, a cancel, or a power failure. Its job is to repair the damage: back out work that was in flight and complete work that was committed.
What emergency restart does
- The recovery manager reads the system log backwards and forwards to find every unit of work that was active when CICS died.
- Units of work that had not reached syncpoint are backed out, so their partial changes disappear.
- Units of work that had committed are completed (forward recovery), so no committed change is lost.
- CICS also resynchronizes with connected systems - remote CICS regions, DB2 and MQ - to settle in-doubt units of work.
How it is triggered
- It is automatic: restart the region after an abnormal end with START=AUTO and CICS detects the abnormal end and runs emergency restart.
- You can also force it explicitly:-//CICSSTEP EXEC PGM=DFHSIP,
// PARM='SI,START=AUTO,APPLID=CICSPROD' - During the restart, operators may see WTOR messages asking for decisions - for example, whether to proceed when a journal is unavailable. Reply carefully.
After emergency restart
- Check for in-doubt or failed units of work with the operator command CEMT INQUIRE UOW.
- Any UOW shown as FAILED or INDOUBT needs attention - resynchronize it or force it after verifying the data.
- Verify that critical files are open and enabled, and that application teams confirm data integrity before users return.
