Sunday, September 20, 2026

𝗖𝗢𝗕𝗢𝗟 𝗣𝗿𝗼𝗴𝗿𝗮𝗺 𝗔𝗯𝗲𝗻𝗱𝘀 𝘄𝗶𝘁𝗵 𝗦𝗕𝟯𝟳 𝘄𝗵𝗶𝗹𝗲 𝗖𝗟𝗢𝗦𝗶𝗻𝗴 𝘁𝗵𝗲 𝗳𝗶𝗹𝗲

A COBOL program successfully executes all its WRITE statements and reaches the CLOSE statement. Then, unexpectedly, the program issues an SB37 space abend.
 
Why would an out-of-space condition become visible during CLOSE rather than during WRITE?
 
Let's investigate.
 
The Test Program
 
Consider this simple COBOL program:
 
        IDENTIFICATION DIVISION.                                    
        PROGRAM-ID. TESTPGM.                                        
        ENVIRONMENT DIVISION.                                       
        CONFIGURATION SECTION.                                      
        SPECIAL-NAMES.                                              
        INPUT-OUTPUT SECTION.                                       
        FILE-CONTROL.                                               
            SELECT OUTPUT-FILE       ASSIGN TO OUTFILE.             
        DATA DIVISION.                                              
        FILE SECTION.                                               
        FD  OUTPUT-FILE.                                            
        01  OUTPUT-REC    PIC X(30).                                
        WORKING-STORAGE SECTION.                                    
        01 WS-GRP.                                                  
           05 FILLER         PIC X(05) VALUE 'WS-I:'.               
           05 WS-I           PIC S9(8) COMP.                        
        01 WS-VAR            PIC  9(8).                             
        PROCEDURE DIVISION.                                         
           ACCEPT WS-VAR                                            
           DISPLAY WS-VAR                                           
           OPEN OUTPUT OUTPUT-FILE                                  
           PERFORM VARYING WS-I FROM 1 BY 1 UNTIL WS-I > WS-VAR     
             WRITE OUTPUT-REC    
           END-PERFORM           
           DISPLAY 'WS-I ' WS-I  
           CLOSE OUTPUT-FILE     
           GOBACK.               
 
The program simply accepts a record count and writes that many 30-byte records.
 
The corresponding JCL is:
 
//GO1    EXEC  PGM=TESTPGM
//SYSOUT   DD  SYSOUT=*                                          
//OUTFILE  DD  DSN=USERID.TESTFIL,
//          LRECL=30,RECFM=FB,VOL=(,,,1),                       
//          DISP=(,CATLG,CATLG),
//          SPACE=(TRK,(1,1),RLSE),BUFNO=5   
//*                                                              
//SYSIN    DD *                                                  
00034520                                                         
//*                                                              
 
So the program attempts to write: 34,520 records
 
The important file allocation characteristics are:
 
LRECL = 30
BLKSIZE = 27,990
SPACE = TRK,(1,1) = 1 Primary + 15 secondary tracks = total 16 tracks
BUFNO = 5
 
First, Let's Look at the Blocks
 
With fixed-block records:

 
Records per block = BLKSIZE / LRECL = 27,990 / 30 = 933 records
 
Therefore, one physical block can contain 933 logical COBOL records.
 
For 34,520 records:  34,520 / 933 = 36 remainder 932
 
So the output consists of:
  • 36 complete blocks
  • 1 partial block containing 932 records
 
Therefore the output dataset needs space for 37 blocks.
 
A track can accommodate two blocks, and each block holds 933 records.
 
With 16 tracks available(based on SPACE parameter), the total number of records that can be successfully written to output file is:
 
16 tracks × 2 blocks per track × 933 records per block = 29,856 records
 
So, the output data set can accommodate exactly 29,856 records before the space-related failure occurs.
 
If we translate 29,856 records to number of blocks, 29856/933 = 32 blocks.
 
But program is trying to write 37 blocks of records.
 
Now Look at What the Program Reports
 
The SYSOUT contains:
 
00034520   
WS-I 00034521                            
IGZ0034W The file with system-name OUTFILE could not be extended.  Secondary extents were not specified or were not available.  The last WRITE was at offset X'E23156CC' in program TESTPGM.                                  
 
The interesting part is:  “WS-I 00034521”
 
This tells us that the PERFORM loop finished.  In other words, COBOL executed all 34,520 WRITE statements.
 
Yet, after the job abended, the output data set contained only 29,856 records.
 
This means that the remaining records were in BUFFER.
 
CLOSE Is More Than "Close the File"
 
Although all the WRITE statements have completed, some records are still held in the output buffers.
 
During CLOSE, these buffered records are written to disk.
 
Since the data set cannot be extended to accommodate them, the program encounters an SB37 abend while closing the file.

Tuesday, September 15, 2026

Don't Trust Connect Time in SMF Type 30 for FICON Environments

I received the following tip from Enterprise Performance Strategies. 

If you're using SMF Type 30 connect time to gauge I/O performance, be careful - in FICON environments, that number is essentially meaningless.
 
Per IBM's documentation on the SMF 30 record's IO_CONNECT_SEC field: "The value of RqsvAIC for the FICON® channel utilization cannot be calculated. In this case, the system adjusts the connect time for FICON DASD to be 1 millisecond per request." (source https://www.ibm.com/docs/en/z-logdata-analytics/5.1.0?topic=data-smf-30-v2-record-type)

Here's why that matters: connect time has historically represented actual data transfer time. But when FICON introduced concurrent active I/Os per channel, that measurement got too complicated to capture accurately - so the system just hard-codes it to a flat 1ms per request instead.
 
Years ago, when average I/O response times were around 3ms, that 1ms floor was a reasonable stand-in. Today, with response times often well under 1ms (some experts suggest 0.1–0.2ms is closer to reality) that artificial 1ms number is now larger than many actual total response times. It throws off any comparison against more granular records like Type 42 (datasets) or Type 101 (DB2).
 
Bottom line: Don't rely on Type 30 connect time in FICON environments to judge batch I/O performance - use Type 42, Type 101, or other record types that capture true connect time instead.

Sunday, September 13, 2026

𝗤𝘂𝗶𝗰𝗸𝗹𝘆 𝗘𝘀𝘁𝗶𝗺𝗮𝘁𝗶𝗻𝗴 𝘁𝗵𝗲 𝗦𝗶𝘇𝗲 𝗼𝗳 𝗬𝗼𝘂𝗿 𝗣𝗿𝗼𝗱𝘂𝗰𝘁𝗶𝗼𝗻 𝗗𝗕𝟮 𝗗𝗮𝘁𝗮𝗯𝗮𝘀𝗲𝘀

If all your production DB2 table spaces and index spaces are allocated within a dedicated SMS storage group, determining the approximate size of your production databases becomes a straightforward task.
 
A quick way to estimate the size of your production DB2 databases is by using ISMF (Interactive Storage Management Facility). Navigate to the Storage Group option and examine the storage group dedicated to your production DB2 environment. Pay particular attention to the following metrics:
 
  • Total Space
  • Free Space 

By subtracting the Free Space from the Total Space, you can calculate the used space within the storage group.

Used Space = Total Space − Free Space
 
The resulting value provides a good estimate of the storage currently consumed by your production DB2 databases.

Saturday, September 12, 2026

𝗧𝗵𝗲 𝗠𝘆𝘀𝘁𝗲𝗿𝘆 𝗼𝗳 𝘁𝗵𝗲 𝟵𝟵% 𝗦𝘁𝗼𝗿𝗮𝗴𝗲 𝗚𝗿𝗼𝘂𝗽 𝗙𝘂𝗹𝗹 𝗔𝗹𝗲𝗿𝘁: 𝗔 𝗜𝗻𝘃𝗲𝘀𝘁𝗶𝗴𝗮𝘁𝗶𝗼𝗻 𝗦𝘁𝗼𝗿𝘆

One early morning, an alert appeared that immediately caught the attention of the infrastructure team: A production database storage group had reached 99% utilization.

For any infrastructure professional, utilization at that level is cause for concern. When storage approaches full capacity, new DB2 dataset allocations can fail, utilities may abend, and application stability can be put at risk.
 
What made this incident particularly intriguing was the speed at which it unfolded. Within a very short period, storage consumption had surged dramatically, triggering automated alerts warning of insufficient space.
 
The obvious question was:  What Consumed So Much Storage So Quickly?
 
Following the Trail
 
A closer examination of the system logs revealed the first clue. The storage management subsystem was reporting that free capacity was rapidly diminishing, and soon afterward, DB2 began issuing messages that certain datasets could not be allocated because the storage group had run out of available space.
 
At first glance, it appeared that the production environment was genuinely running out of space.
 
Then something unexpected happened.
 
Just minutes later, storage utilization dropped significantly without any administrative intervention. The apparent storage crisis vanished as quickly as it had appeared.
 

Now there were two mysteries to solve:

  1. What caused the sudden spike in storage consumption?
  2. Why did utilization return to normal so quickly?

Correlating the Timeline
 
The breakthrough came when the storage alerts were correlated with active database maintenance activity.
 
At the exact time the utilization spike occurred, a DB2 utility job was performing an online table reorganization (REORG) on one of the largest tables in the environment.
 
Online REORG is designed to maintain application availability while reorganizing data. To achieve this, DB2 creates temporary shadow datasets that exist alongside the original table data throughout the reorganization process.
 
These shadow datasets consumes nearly as much storage as the original table itself.
 
As the utility progressed, storage consumption rose rapidly until the storage group was effectively exhausted. Eventually, the REORG utility failed because additional space could no longer be allocated.
 
The final piece of the puzzle appeared moments later.
 
When the failed utility terminated, DB2 automatically cleaned up the temporary shadow datasets it had created. As those datasets were deleted, storage utilization immediately dropped back to normal levels.
 
Case Closed
 
What initially appeared to be a critical and unexplained storage shortage was actually the side effect of an online REORG operating on a very large table.
 
The temporary shadow datasets created by the utility drove storage utilization to 99%, triggering alerts and allocation failures. Once the REORG failed and cleaned up its shadow files, the consumed space was released, causing utilization to fall back to normal almost instantly. 

Sunday, September 6, 2026

𝗧𝗿𝗮𝗻𝘀𝗳𝗼𝗿𝗺𝗶𝗻𝗴 𝗮 𝗖𝗢𝗕𝗢𝗟 𝗔𝗿𝗿𝗮𝘆 𝗶𝗻𝘁𝗼 𝗮 𝗧𝗲𝗺𝗽𝗼𝗿𝗮𝗿𝘆 𝗧𝗮𝗯𝗹𝗲 𝗮𝗻𝗱 𝗝𝗼𝗶𝗻𝗶𝗻𝗴 𝘄𝗶𝘁𝗵 𝗮 𝗗𝗕𝟮 𝘁𝗮𝗯𝗹𝗲

It is not uncommon for COBOL programs to maintain a list of values in a WORKING STORAGE table and use those values to search for matching rows in a DB2 table. Consider the following working-storage structure:
 
01 WS-DIAG-CODE-TABLE.
   05 WS-DIAG-CODE-COUNT                   PIC S9(04) COMP.
   05 WS-DIAG-CODE-LIST.
      10 WS-DIAG-CODE OCCURS 1 TO 100 TIMES
          DEPENDING ON WS-DIAG-CODE-COUNT    PIC X(05).
 
Traditional Approach
 
The typical implementation is to perform a loop through each occurrence of WS-DIAG-CODE and execute a DB2 query to fetch matching rows.
 
PERFORM VARYING IDX FROM 1 BY 1
   UNTIL IDX > WS-DIAG-CODE-COUNT
   EXEC SQL
     SELECT ...
       FROM TEST_TABLE
     WHERE DIAG_CODE = :WS-DIAG-CODE(IDX)
   END-EXEC
END-PERFORM
 
While this approach is straightforward and easy to understand, it can become a significant performance bottleneck. Each iteration results in a separate database access, increasing CPU consumption and elapsed time.
 
A More Efficient Alternative
 
A better approach is to convert the COBOL table into a temporary result set within SQL and allow DB2 to process all diagnosis codes in a single query.
 
This can be achieved using a recursive Common Table Expression (CTE). The recursive CTE transforms the contents of the COBOL array into a relational structure that can be joined directly with the target DB2 table.
 
WITH TEMP (IDX, DIAG_CODE) AS
(SELECT 1, LEFT(:WS-DIAG-CODE-LIST, 5)
   FROM SYSIBM.SYSDUMMY1
UNION ALL
SELECT IDX + 1, SUBSTR(:WS-DIAG-CODE-LIST, (IDX * 5) + 1, 5)
  FROM TEMP
WHERE IDX < :WS-DIAG-CODE-COUNT),
SELECT ...
FROM TEST_TABLE A, TEMP T
WHERE A.DIAG_CODE = T.DIAG_CODE;
 
Why This Approach Performs Better
 
By leveraging a recursive CTE:
  • Only one SQL statement is executed.
  • DB2 can optimize the join operation internally.
  • Database round trips are eliminated.
  • CPU and elapsed time are significantly reduced, especially when the list contains many values.
  • The solution scales better as the number of diagnosis codes grows.
 
Key Takeaway: Whenever you find yourself executing the same SQL repeatedly for values stored in a COBOL OCCURS table, consider transforming the data into a temporary table using a recursive CTE and let DB2 perform the matching in one pass. This is a classic example of replacing row-by-row processing with high-performance set-based SQL.