pca(1)
pca - Power C Analyzer
Showing IRIX 6.5.15. Added in IRIX 6.5.5.
NAME pca - Power C Analyzer SYNOPSIS /usr/lib/pca [ options ] ... /usr/lib64/cmplrs/pca [ options ] ... IMPLEMENTATION IRIX systems (-o32 ABI only) DESCRIPTION For -n32 and -64 ABIs, pca has been replaced by the -apo option. pca is a source-to-source optimizing C preprocessor that discovers parallelism in C code. pca concurrentizes C by restructuring certain portions of code and then adding parallel programming directives where possible. The amount of code that is concurrentized varies depending on the value of certain pca command line options, directives, and assertions. For more information, refer to the IRIS Power C User's Guide. pca is usually invoked as an option to the cc(1) command, although it can be run separately. When pca is used as part of a cc compilation, the pca options must be passed via the -WK mechanism. See the cc(1) man page for details of the -W option. pca requires a single input source file, which can be specified using the -i option. Specifying a filename as a command-line argument is also sufficient, as any argument not recognized as an option is treated as an input filename. pca can produce two types of output. The first type is the concurrentized C source code, and the second type is the annotated listing, whose contents vary depending on the value of the -lo command line option. Filenames for these two types of output can be set using the -cmp and -l options, respectively. If the -cmp option is not specified, the output code will be sent to standard output. If no -l option is specified, and some type of listing information is requested using the -lo option, then this listing information will also be sent to standard output. pca accepts the following command line options: -arl=<integer> Long name: -address_resolution_level=<integer> Default value: -address_resolution_level=1 The -arl option controls the assumptions pca makes about memory aliases. The integer value, which ranges from 0 (no assumptions) to 4 (assume everything), specifies the level of assumptions that pca is to make about the behavior of the code. Each level is cumulative (i.e., level 4 also makes all the assumptions of the lower levels). The following are the assumption level definitions: 0 No assumptions are made. 1 Assumes that no pointer self references. 2 Assumes that function arguments are distinct from each other. 3 Assumes that local pointers and arrays are distinct from global pointers and arrays. 4 Assumes that all pointers and arrays are distinct from each other. -arclm=<integer> Long name: -arclimit=<integer> Default value: -arclimit=5000 The -arclimit option sets the size of the dependence arc data structure that pca uses to perform data dependence analysis. This data structure is dynamically allocated on a loop nest by loop nest basis. By default, this data structure is allocated with a size = max (# of statements * 4, -arclimit value). If a loop contains too many dependence relationships and cannot be represented in the dependence data structure, pca will surrender optimization of the loop. You may use the -arclimit option to increase the size of the data structure to enable pca to perform more optimizations. (Most users do NOT need to change this value.) -chl=<integer> Long name: -cacheline=<integer> Default value: -cacheline=64 The -cacheline option informs pca of the width of the memory channel (in bytes) between cache and main memory. -chs=<integer> Long name: -cachesize=<integer> Default value: -cachesize=64 The -cachesize option informs pca of the size (in kilobytes) of the cache memory. -cplc=<integer>[,<integer>] Long name: -cache_prefetch_line_count=<integer>[,<integer>] Default value: -cache_prefetch_line_count=0 The -cplc option informs pca of the number of additional cache lines that are fetched during a cache miss. The first number corresponds to the primary cache, and the second number, if specified, corresponds to the secondary cache. -[n]cmp[=<file>] Long name: -[no]cmp[=<file>] Default value: standard output The -cmp (compilable) option instructs pca to place the optimized C program in a specified transformed program file. If -cmp=<file> is specified, the transformed C is written to that file. If -cmp is specified, the transformed code is written to filename.m, where filename is the input file name with any trailing .c removed. If no -cmp option is specified, the transformed code is written to standard output. To disable generation of the C output file, specify -nocmp on the command line. -[n]conc Long name: -[no]concurrentize Default value: -concurrentize This command line option directs pca to perform concurrentization, while -noconc directs pca to perform only scalar optimizations. -[n]cp=<list> Long name: -[no]cmpoptions=<list> Default value: -nocmpoptions The -cmpoptions option specifies additional information for inclusion in the transformed code file. The following option can be specified: i Insert line numbers that reference the original source. -dollar No short name. Default value: off The -dollar option allows dollar signs to be used in identifiers under both Kernighan and Ritchie and ANSI C. -dpr=<integer> Long name: -dpregisters=<integer> Default value: -dpregisters=12 The -dpregisters option specifies the number of double precision floating point registers each processor has available for expression evaluation. -eiifg=<integer> Long name: -each_invariant_if_growth=<integer> Default value: -each_invariant_if_growth=20 The -each_invariant_if_growth option controls the rewriting of ifs nested within loops. The value of the option (which must be in the range 0 to 100) specifies a limit on the number of executable statements in the nested if. If the number of statements in the if exceeds this limit, then pca will not rewrite the code. If there are fewer statements, then pca will interchange the loop and if statements for improved execution speed. -float No short name. Default value: off In Kernighan and Ritchie C, all variables declared as type float are promoted to type double by default. The -float option prevents this promotion to double. This option is ignored under -syntax=a, because ANSI C does not promote variables of type float to double. -fpr=<integer> Long name: -fpregisters=<integer> Default value: -fpregisters=12 The -fpregisters option specifies the number of single precision floating point registers each processor has available for expression evaluation. -[n]fuse Long name: -[no]fuse Default value: -fuse This command line option enables loop fusion, a conventional compiler optimization that transforms two adjacent loops into a single loop. Setting -scalaropt to at least 2, -optimize=5 is also required to enable loop fusion. -heap=<list> Long name: -heaplimit=<integer> Default value: -heaplimit=100 The -heaplimit option specifies the maximum size in megabytes to which the memory heap is allowed to grow to when pca processes a source file. If this limit is reached, pca will stop processing the source code. -hli=<list> Long name: -hoist_loop_invariants=<integer> Default value: -hoist_loop_invariants=1 This option controls code hoisting of loop-invariant expressions from loops. Note that this switch is independent of the switches that control the floating of invariant-IFs out of loops, -each_invariant_if_growth and -max_invariant_if_growth. This option accepts the following settings: 0 Turns off the hoisting of invariant code from loops. 1 Floats all loop invariant expressions that are not under the control of an IF-structure within the given loop nest. This is the default setting. 2 The same behavior as level 1. 3 Floats all loop-invariant expressions from loops. If there is invariant code that is protected by an IF- structure and the hoisting value is less than 3, then pca will generate a message to that effect in your output listing. -i=<file> Long name: -input=<file> Default value: none The -i option instructs pca to read its input code from the specified file. -inl=<list> Long name: -inline=<list> Default value: off -inlc=<list> Long name: -inline_and_copy=<list> Default value: off -ipa=<list> Long name: -ipa=<list> Default value: off The -inline, -inline_and_copy, and -ipa options provide pca with a list of routines to analyze. If the option is specified without an argument list, pca will try to inline/analyze all the called functions in the inlining universe (specified by the -inline_from_files, -inline_from_libraries, -ipa_from_files, or -ipa_from_libraries options). If a list of names is included, for example, -inline=mkcoef,yval, then just the routines named will be inlined/analyzed. The -inline and -ipa command line options are overridden by the #pragma inline and #pragma ipa directives. The -inline_and_copy option functions like the -inline option, except that if all references to a function are inlined, the text of the routine is not optimized, but is copied unchanged to the transformed code file. This is intended for use when inlining routines from the same file as the call, and has no special effect when the routines being inlined are being taken from a library or another source file. When a subprogram has been inlined everywhere it is used, leaving it unoptimized saves compilation time. When a program involves multiple source files, the unoptimized routine will still be available in case one of the other source files contains a reference to it, so no errors will result. NOTE: the inline_and_copy algorithm assumes that all calls and references to the routine precede it in the source file. If the routine is referenced after the text of the routine, and that particular call site cannot be inlined, the unoptimized version of the routine will be invoked. -incr=<file> Long name: -inline_create=<file> Default value: off -ipacr=<file> Long name: -ipa_create=<file> Default value: off These options instruct pca to build a library file containing partially-analyzed routines for later inlining/analyzing. The library created is used with the -inline_from_libraries or -ipa_from_libraries option. Libraries created with -inline_create can be used with either inlining or interprocedural analysis, since they contain essentially complete descriptions of the functions included. Libraries created with -ipa_create can be used only with interprocedural analysis, since they do not have the complete text of the functions, just the data relationships information. Any filename can be used for the library name. An extension .klib is preferred, for maximum compatibility with the -inline_from_libraries and -ipa_from_libraries options. -ind=<integer> Long name: -inline_depth=<integer> Default value: -inline_depth=2 This option sets the maximum level of subprogram nesting that pca will attempt to inline. Higher values instruct pca to trace function calls further and thus produce more inlined code. The #pragma [no]inline directive, when enabled, is not affected by the -inline_depth restrictions. The argument values and their meanings are: -1 Inline only routines that do not contain procedure or function calls. 0 Use the default value (2). 1-10 Inline routines to this depth. -inff=<list> Long name: -inline_from_files=<list> Default value: current source file -infl=<list> Long name: -inline_from_libraries=<list> Default value: off -ipaff=<list> Long name: -ipa_from_files=<list> Default value: current source file -ipafl=<list> Long name: -ipa_from_libraries=<list> Default value: off These options provide pca with the locations of functions available for inlining/interprocedural analysis. (The total set of available functions is called the inlining (or IPA) universe.) The -inline_from_files and -ipa_from files options take the names of source files and directories containing source files. Including a directory, for example, -ipaff=/usr/ipalib, is equivalent to the UNIX notation /usr/ipalib/*.c. (Do not use shell wild card characters in the list of files and directories.) Note that the specified files must have already been pre-processed by cpp or acpp (if necessary). The -inline_from_libraries and -ipa_from_libraries options take the names of libraries created with the -inline_create and -ipa_create options and directories containing such libraries. In directories, the pca libraries are identified by the extension .klib. Multiple files/libraries or directories may be specified in one option, separated by commas. Multiple occurrences of these options may be specified on the command line. -inll=<integer> Long name: -inline_looplevel=<integer> Default value: -inll=2 -ipall=<integer> Long name: -ipa_looplevel=<integer> Default value: -ipall=2 The -inline_looplevel and -ipa_looplevel options enable the user to limit inlining to just functions which are referenced in nested loops, where the effects of reduced function call overhead or enhanced optimizations will be multiplied. The parameter is defined from the most deeply nested function reference. -inll=1 restricts inlining to functions referenced in the deepest loop nest. -inll=3 restricts inlining to those routines referenced at the three deepest levels. The for loop nest level of each function reference is included in the optional calling tree section of the listing files. The #pragma [no]inline and #pragma [no]ipa directives, when enabled, are not affected by the -inline_looplevel and -ipa_looplevel restrictions. -inm Long name: -inline_manual Default value: off -ipam Long name: -ipa_manual Default value: off These options instruct pca to recognize the #pragma [no]inline and #pragma [no]ipa directives. This allows manual control over which functions are inlined/analyzed at which call sites. The default is to ignore these #pragma directives. When -inline_man or -ipa_man is included on the command line, the #pragma directives are enabled. Because #pragma [no]inline and #pragma [no]ipa are not affected by the -inline_looplevel and -ipa_looplevel command line options, they can be used with command line control to select functions or call sites that the regular selection algorithm would reject. -lm=<integer> Long name: -limit=<integer> Default value: -limit=5000 To reduce the compile time, pca estimates how long it takes to analyze each loop nest construct. If a loop is too deeply nested, pca ignores the outer loop and recursively visits the inner loops. The loop nest limit can be used to control what pca considers too deeply nested. For more information, refer to the IRIS Power C User's Guide. -ln=<integer> Long name: -lines=<integer> Default value: -lines=55 This option enables the pca listing to be paginated for printing. The number of lines per page on the listing may be changed by using the -lines option. The -lines=0 option directs pca to paginate only at subroutine boundaries. -[n]l[=<file>] Long name: -[no]list[=<file>] Default value: standard output If some type of listing information is requested using the -lo option, the -list option specifies the destination of that listing. If -list=<file> is specified, the listing is written to the specified file. If -list is specified without a filename, the listing file is written to filename.lst, where filename is the input file name with any trailing .c removed. If no -list option is specified, any listing information requested is written to the standard output. To disable generation of any listing information, enter -nolist on the command line. -lo=<list> Long name: -listoptions=<list> Default value: no listing The -listoptions option tells pca what information to include in the listing file. -listoptions can enable the following options to be included in the listing: c Calling tree. k Print pca options active within the program unit. l Loop-by-loop optimization table. n Program unit names as processed (written to error file). p Compilation performance statistics. s Summary of the optimizations performed. -lw=<integer> Long name: -listingwidth=<integer> Default value: -listingwidth=80 This option sets the maximum line length for the listing file. This setting affects only the format of the loop summary table (-lo=l) and the options table (-lo=k). The only permissible values for -listingwidth are 80 and 132. -[n]ma=<list> Long name: -[no]machine=<list> Default value: -machine=s The -machine option is list-valued with three choices which may be set. The possible choices are: n Tells pca to concurrentize loops which generate non- stride-1 memory references. o Tells pca to only concurrentize outer loops. This capability is available to prevent concurrentization on applications that have small inner loop bounds, thereby reducing overhead costs. pca determines concurrentization based on the overhead to benefit ratio. When the loop bounds are unknown at compile time, pca may generate concurrent code for innermost loops, a practice that may be inefficient for the actual loop bounds. s Tells pca to concurrentize for loops which generate stride-1 memory references. The s and n choices may not be set simultaneously. -miifg=<integer> Long name: -max_invariant_if_growth=<integer> Default value: -max_invariant_if_growth=500 The -max_invariant_if_growth option controls the rewriting of if statements nested within loops. The value of the option (which must be in the range 0 to 1000) specifies a limit on the total number of additional executable statements which can be generated by this optimization. If the total number of extra statements generated exceeds this limit, then pca will stop rewriting the code. If there are fewer statements, then pca will continue to interchange the loop and if statements for improved execution speed. -mc=<integer> Long name: -minconcurrent=<integer> Default value: -minconcurrent=1000 The -minconcurrent option sets the minimum level of work which must be done by a loop in order for it to be a candidate for concurrentization. Loops that contain less work than -minconcurrent will not be concurrentized. Setting -minconcurrent=0 effectively allows all loops (if all other necessary conditions are met) to be concurrentized. -[n]namepart[=<integer>,<integer>] Long name: -[no]namepartitioning[=<integer>,<integer>] Default value: -nonamepartitioning This option looks at distinct array names and limits the number of arrays that that are referenced within a loop in order to avoid cache conflicts among them. For example, -namepart can be used to break a loop that contains references to arrays a and b into two loops, one referencing array a and the other referencing array b. The two integer arguments specify the preferred minimum and maximum numbers of distinct arrays referenced in each distributed loop. If they are not specified, they default to 2 and 6. The -nnamepart option turns off name partitioning. -scalaropt must be set to at least 3 for namepartitioning to work. -o=<integer> Long name: -optimize=<integer> Default value: -optimize=5 The -optimize option sets the optimization level. Each optimization level is cumulative (i.e., level 5 performs everything up to and including this level). The meaning of the optimization levels are as follows: 0 No optimization is done. 1 Only simple optimizations are done. Induction variable recognition is enabled. 2 Recognizes reductions in concurrent loops. Performs lifetime analysis to determine when last-value assignment of scalars is necessary. Performs more powerful data dependence tests to find parallelism. 3 Special techniques are used to break data dependence cycles that otherwise prevent concurrentization. Triangular loops are recognized and loop interchanging will be attempted to improve memory referencing. Special case data dependence tests are used. Special index sets, called wrap-around variables, are also recognized. 4 Two versions of a loop are generated, if necessary, to break a data dependence arc. Exact data dependence tests are used to allow more parallelism to be discovered. 5 Array expansion and loop fusion are enabled. -r=<integer> Long name: -roundoff=<integer> Default value: -roundoff=0 The -roundoff option allows the user to specify the change from serial roundoff error that is tolerable. If an arithmetic reduction is accumulated in a different order than in the scalar program, the roundoff error is accumulated differently and the final result may differ from that of the original program's output. Although the difference is usually insignificant, certain restructuring transformations performed by pca must be disabled to obtain exactly the same answers as the scalar program. pca classifies its transformations by the amount of difference in roundoff error that can accumulate so the user can decide what level of roundoff error differences is allowable. The meanings of the switch values are as follows: 0 Allow no roundoff-changing transformations. Arithmetic recurrences and reductions are not concurrentized. Non-arithmetic reductions may still be concurrentized. Thus, the answers do not depend on the number of processors used. 1 Enable expression simplification which might generate overflow or underflow errors differently. Enable simplification of expressions with operands between binary and unary operators. Perform expression simplification which is exposed due to forward substitution. Enable code floating if the -scalaroptimize option is greater than or equal to 1. 2 Allow loop interchanging around arithmetic reductions to be recognized. Perform concurrent reductions with pre-scheduled concurrent loops and local accumulation of reduction results. 3 Recognize real (float) induction variables. Enable sum reductions. Enable memory management optimizations if scalaroptimize=3. -routine=<routine>[,<routine>...] -skip This option pair allows the user to request that pca perform no optimization on the specified routines. The code in these routines will be written unchanged to the output file. -signed No short name. Default value: off By default, a variable declared as char is handled as an unsigned char. The -signed option causes a variable declared as char to be handled as signed char. This option is necessary when porting code from platforms that have a C compiler that defaults char to signed char. -so=<integer> Long name: -scalaroptimize=<integer> Default value: -scalaroptimize=3 The -scalaroptimize option sets the level of dusty-deck and other serial transformations performed. The following are the possible values for <integer>: 0 No scalar optimizations are performed. 1 Only simple scalar optimizations are performed - dead code elimination, global forward substitution and dusty-deck IF transformations. 2 Full range of scalar optimizations are performed - invariant IF's are floated out of loops, loop rerolling/unrolling, array expansion, loop fusion, loop peeling and induction variable recognition. 3 Memory management optimizations are enabled if roundoff=3. 4 Reruns dead code elimination optimization after other optimizations. Sometimes the other optimizations expose more dead code which can be removed. -sasc=<integer> Long name: -setassociativity=<integer> Default value: -setassociativity=1 The -setassociativity option provides information on the mapping of physical addresses in main memory to cache pages. The default value (1) allows a datum in main memory to be placed in any of four places in cache. -spr=<integer> Long name: -spregisters=<integer> Default value: -spregisters=12 The -spregisters option specifies the number of single precision floating point registers each processor has available for expression evaluation. -stdio No short name. Default value: off The -stdio option instructs pca to perform strength reduction on calls to certain functions in the standard I/O library. Programs that use functions such as printf or scanf heavily will generally have improved I/O performance when this switch is used. The -scalaroptimize=3 option is required to enable these transformations. Use of this option may produce code containing calls to functions defined in libkapio.a, a library of optimized low-level I/O routines. As a result, when linking your program using cc(1) or ld(1), you may have to include this library, using the command-line option -lkapio. If you use cc(1) with a -pca option to link your program, this is done automatically. If you explicitly specify that I/O using scanf or fscanf be done in parallel (for example, by using the #pragma concurrent call directive), using -stdio might cause your program to yield incorrect results. To determine if your program might have this problem, search the pca output code (if you are using cc(1), use the -pca keep option to put the output code in a .M file) for the functions __KIOrd, __KIOrf, and __KIOri. If you find these functions (from optimized calls to scanf or fscanf on doubles, floats, and integers, respectively), the I/O read operations involved will not be semaphored properly. This is only a problem if they occur in a parallel region, and operate on files that other threads may also be operating on concurrently. If you do no explicit specification of parallel regions, use of the -stdio option is safe. -su=<list> Long name: -suppress=<list> Default value: no message suppression The -suppress option tells pca to disable the printing of individual classes of messages. These message classes range from syntax warning and error messages to messages about the optimizations performed. The following codes can be specified with the -suppress option: d Data dependence messages. e Syntax error messages. i Informational messages. n "Not optimized" messages. q Questions. s Standardized messages. w Syntax warning messages. -sy=<list> Long name: -syntax=<list> Default value: -syntax=a The -syntax option tells pca to check for compliance with certain syntactic rules. The -syntax options are: a Check for compliance with the SGI-extended Ansi C standard. k Check for compliance with the SGI-enhanced Kernighan & Richie C specification. -ur=<integer> Long name: -unroll=<integer> Default value: -unroll=4 -ur2=<integer> Long name: -unroll2=<integer> Default value: -unroll2=100 -ur3=<integer> Long name: -unroll3=<integer> Default value: -unroll3=1 The -unroll, -unroll2, and -unroll3 options control how pca unrolls inner loops. The value of the -scalaroptimize option must be at least 2 in order for unrolling to be performed. The unrolling that pca performs is also subject to the following constraints imposed by these three options: * The number of times a loop is unrolled will not be larger than the value given by the -unroll option. If the value given is zero, the default value is used instead. * The work done by the loop body is estimated by counting operands and operators. This estimate of the amount of work done by the unrolled loop body must be less than the value of the -unroll2 option and greater than the value of the -unroll3 option. The reason for the -unroll and -unroll2 options is that beyond a certain point, further unrolling of a loop, while continuing to increase code size, brings diminishing returns in performance since the work being done in the loop body is already large compared to the overhead associated with a loop iteration, which unrolling is attempting to amortize. The reason for the -unroll3 option is to prevent unrolling of very small loops which might be better unrolled or software-pipelined by the C compiler. Since these compiler optimizations are not on by default, the default value for this option is 1. If you would like to take advantage of these C compiler capabilities, you should set this option to a higher value. If the constraints imposed by the option values cannot be met for a given loop, then it will not be unrolled. If they can be met, then the largest degree of unrolling that satisfies the constraints will be used. -volatile No short name. Default value: off The -volatile option indicates that all variables are implicitly volatile. Use of this option severely limits the optimization that can be done. -[n]64 Long name: -[no]64 Default value: off This option tells pca that the code is being compiled for a 64-bit machine. This affects the size of certain C data types. FILES /usr/lib/pca The pca program /usr/lib64/cmplrs/pca (32-bit and 64-bit versions) file.c C source file file.m C transformed file file.lst listing file /usr/lib/libkapio.a Optimized I/O library (for -stdio option) /usr/lib64/mips3/libkapio.a (mips1, mips3, and mips4 versions) /usr/lib64/mips4/libkapio.a SEE ALSO cc(1), cpp(1), mpc(1) IRIS Power C User's Guide IRIS Power C Quick Reference Parallel Programming on Silicon Graphics Computer Systems Practical Parallel Programming by Dr. Barr Bauer, Academic Press, 1991. Introduction to Parallel Programming by Steven Brawer, Academic Press, 1989. This man page is available only online.