Module 13: CICS Recovery, Restart and Production Support
Production Troubleshooting
Troubleshooting is a discipline, not guesswork. Follow the same method every time - gather evidence, form a theory, test it, fix it - and most CICS production problems resolve quickly.
The troubleshooting method
- Reproduce or characterize: who is affected, when did it start, what changed just before it started.
- Read the evidence first: the MSGUSR log, the console log, the abend code and any dumps - before touching anything.
- Use CEMT to inspect live state: INQUIRE TASK, INQUIRE FILE, INQUIRE UOW and INQUIRE TRANSACTION show what CICS is doing right now.
- Change one thing at a time, and verify the fix with the users who reported the problem.
Key diagnostic sources
- MSGUSR and the job log - every CICS message, abend code and WTOR reply is recorded here.
- Transaction and system dumps - formatted with IPCS, they show exactly where a program failed.
- CICS statistics (DFHSTUP) - task rates, file I/O, storage use and waits, invaluable for performance issues.
- Auxiliary trace - a last resort that records CICS internal calls; use it briefly because of the volume it generates.
Printing CICS statistics
- Run the DFHSTUP utility against the statistics dataset to get a formatted report:-//STATJOB JOB (ACCT),'CICS STATS',CLASS=A,MSGCLASS=X
//STATS EXEC PGM=DFHSTUP
//SYSUT1 DD DSN=PROD.CICS.STATSDSN,DISP=SHR
//SYSUT2 DD SYSOUT=* - Review the report for rising file I/O times, storage pressure and task queueing before they become outages.
- Keep statistics history so you can compare 'normal' days with bad days - trends beat single snapshots.
Escalation and change control
- If the fix is not obvious within your timebox, escalate with the evidence gathered - abend code, messages, dumps, CEMT output.
- Every production change (NEWCOPY, SET commands, definition installs) should be logged so the next problem has a history to check.
- After a major incident, write down what happened and what fixed it - that note is tomorrow's five-minute fix.
