intro_shmem(3)
intro_shmem - Introduction to logically shared memory access routines
As shipped in IRIX 6.5.19. Added in IRIX 6.5.19.
NAME intro_shmem - Introduction to logically shared memory access routines DESCRIPTION The logically shared, distributed memory access (SHMEM) routines provide low-latency, high-bandwidth communication for use in highly parallelized scalable programs. The SHMEM routines are data passing library routines similar to message passing library routines. They can be used as an alternative to message passing routines such as Message Passing Interface (MPI) or Parallel Virtual Machine (PVM). Like the message passing routines, the SHMEM routines pass data between cooperating parallel processes. SHMEM routines can be used in programs that perform computations in separate address spaces and that explicitly pass data to and from different processes in the program. These processes are also called processing elements (PEs). The SHMEM routines minimize the overhead associated with data passing requests, maximize bandwidth, and minimize data latency. Data latency is the period of time that starts when a PE initiates a transfer of data and ends when a PE can use the data. SHMEM routines support remote data transfer through put operations, which transfer data to a different PE, and get operations, which transfer data from a different PE. Other operations supported are work-shared broadcast and reduction, barrier synchronization, and atomic memory operations. An atomic memory operation is an atomic read-and-update operation, such as a fetch-and-increment, on a remote or local data object. The value read is guaranteed to be the value of the data object just prior to the update. SHMEM Routines The following SHMEM-related routines enhance the portabiliy of SHMEM programs across platforms. * PE queries: C/C++ only: _num_pes(3I), _my_pe(3I) Fortran only: NUM_PES(3I), MY_PE(3I) * Block data put routines: C/C++ and Fortran: shmem_put32, shmem_put64, shmem_put128 C/C++ only: shmem_double_put, shmem_float_put, shmem_int_put, shmem_long_put, shmem_short_put Fortran only: shmem_complex_put, shmem_integer_put, shmem_logical_put, shmem_real_put * Block data get routines: C/C++ and Fortran: shmem_get32, shmem_get64, shmem_get128 C/C++ only: shmem_double_get, shmem_float_get, shmem_int_get, shmem_long_get, shmem_short_get Fortran only: shmem_complex_get, shmem_integer_get, shmem_logical_get, shmem_real_get * Strided put routines: C/C++ and Fortran: shmem_iput32, shmem_iput64, shmem_iput128 C/C++ only: shmem_double_iput, shmem_float_iput, shmem_int_iput, shmem_long_iput, shmem_short_iput Fortran only: shmem_complex_iput, shmem_integer_iput, shmem_logical_iput, shmem_real_iput * Strided get routines: C/C++ and Fortran: shmem_iget32, shmem_iget64, shmem_iget128 C/C++ only: shmem_double_iget, shmem_float_iget, shmem_int_iget, shmem_long_iget, shmem_short_iget Fortran only: shmem_complex_iget, shmem_integer_iget, shmem_logical_iget, shmem_real_iget * Point-to-point synchronization routines: C/C++ only: shmem_int_wait, shmem_int_wait_until, shmem_long_wait, shmem_long_wait_until, shmem_longlong_wait, shmem_longlong_wait_until, shmem_short_wait, shmem_short_wait_until Fortran: shmem_int4_wait, shmem_int4_wait_until, shmem_int8_wait, shmem_int8_wait_until * Barrier synchronization routines: C/C++ and Fortran: shmem_barrier_all, shmem_barrier * Atomic memory fetch-and-operate (fetch-op) routines: C/C++ and Fortran: shmem_swap * Reduction routines: C/C++ only: shmem_int_and_to_all, shmem_long_and_to_all, shmem_longlong_and_to_all, shmem_short_and_to_all, shmem_double_max_to_all, shmem_float_max_to_all, shmem_int_max_to_all, shmem_long_max_to_all, shmem_longlong_max_to_all, shmem_short_max_to_all, shmem_double_min_to_all, *shmem_float_min_to_all, shmem_int_min_to_all, shmem_long_min_to_all, shmem_longlong_min_to_all, shmem_short_min_to_all, shmem_double_sum_to_all, shmem_float_sum_to_all, shmem_int_sum_to_all, shmem_long_sum_to_all, shmem_longlong_sum_to_all, shmem_short_sum_to_all, shmem_double_prod_to_all, shmem_float_prod_to_all, shmem_int_prod_to_all, shmem_long_prod_to_all, shmem_longlong_prod_to_all, shmem_short_prod_to_all, shmem_int_or_to_all, shmem_long_or_to_all, shmem_longlong_or_to_all, shmem_short_or_to_all, shmem_int_xor_to_all shmem_long_xor_to_all shmem_longlong_xor_to_all shmem_short_xor_to_all, Fortran only: shmem_int4_and_to_all, shmem_int8_and_to_all, shmem_real4_max_to_all, shmem_real8_max_to_all, shmem_int4_max_to_all, shmem_int8_max_to_all, shmem_real4_min_to_all, shmem_real8_min_to_all, shmem_int4_min_to_all, shmem_int8_min_to_all, shmem_real4_sum_to_all, shmem_real8_sum_to_all, shmem_int4_sum_to_all, shmem_int8_sum_to_all, shmem_real4_prod_to_all, shmem_real8_prod_to_all, shmem_int4_prod_to_all, shmem_int8_prod_to_all, shmem_int4_or_to_all, shmem_int8_or_to_all, shmem_int4_xor_to_all, shmem_int8_xor_to_all * Broadcast routines: C/C++ and Fortran: shmem_broadcast32, shmem_broadcast64 * Generalized barrier synchronization routine: C/C++ and Fortran: shmem_barrier * Cache management routines: C/C++ and Fortran: shmem_udcflush, shmem_udcflush_line * Byte-granularity block put routines: C/C++ and Fortran: shmem_putmem and shmem_getmem Fortran only: shmem_character_put and shmem_character_get * Collect routines: C/C++ and Fortran: shmem_collect32, shmem_collect64, shmem_fcollect32, shmem_fcollect64 * Atomic memory fetch-and-operate (fetch-op) routines: C/C++ only: shmem_double_swap, shmem_float_swap, shmem_int_cswap, shmem_int_fadd, shmem_int_finc, shmem_int_swap, shmem_long_cswap, shmem_long_fadd, shmem_long_finc, shmem_long_swap, shmem_longlong_cswap, shmem_longlong_fadd, shmem_longlong_finc, shmem_longlong_swap Fortran only: shmem_int4_cswap, shmem_int4_fadd, shmem_int4_finc, shmem_int4_swap, shmem_int8_swap, shmem_real4_swap, shmem_real8_swap, shmem_int8_cswap * Atomic memory operation routines: Fortran only: shmem_int4_add, shmem_int4_inc * Remote memory pointer function: C/C++ and Fortran: shmem_ptr * Reduction routines: C/C++ only: shmem_longdouble_max_to_all, shmem_longdouble_min_to_all, shmem_longdouble_prod_to_all, shmem_longdouble_sum_to_all Fortran only: shmem_real16_max_to_all, shmem_real16_min_to_all, shmem_real16_prod_to_all, shmem_real16_sum_to_all Remotely Accessible Data Objects Typically, target or source arrays that reside on remote processing elements (PEs) are identified by passing the address of the corresponding data object on the local PE. The local existence of a corresponding data object implies that a data object is symmetric as described on this man page. Symmetric data objects passed to SHMEM routines can be arrays or scalars. A symmetric data object is one for which the local and remote addresses have a known relationship. You can use SHMEM routines to access remote symmetric data objects by using the address of the corresponding data object on the local PE. The following data objects are symmetric: * Fortran data objects in common blocks or with the SAVE attribute. These data objects must not be defined in a dynamic shared object (DSO). * Non-stack C and C++ variables. These data objects must not be defined in a DSO. * Fortran arrays allocated with shpalloc(3F) * C and C++ data allocated by shmalloc(3C) SHMEM collective routines that operate on the same data object on multiple PEs require that symmetric data objects be passed. This restriction is for algorithm simplicity and efficiency. These routines define the set of target PEs by the following triplet of arguments: PE_start, logPE_stride, and PE_size. Collective Routines Some SHMEM routines, for example, shmem_broadcast(3) and shmem_float_sum_to_all(3), are classified as collective routines because they distribute work across a set of PEs. They must be called concurrently by all PEs in the active set defined by the PE_start, logPE_stride, PE_size argument triplet. The following man pages describe the SHMEM collective routines: * shmem_and(3) * shmem_barrier(3) * shmem_broadcast(3) * shmem_collect(3) * shmem_max(3) * shmem_min(3) * shmem_or(3) * shmem_prod(3) * shmem_sum(3) * shmem_xor(3) Using the Symmetric Work Array, pSync Multiple pSync arrays are often needed if a particular PE calls a SHMEM collective routine twice without intervening barrier synchronization. Problems would occur if some PEs in the active set for call 2 arrive at call 2 before processing of call 1 is complete by all PEs in the call 1 active set. You can use shmem_barrier() or shmem_barrier_all(3) to perform a barrier synchronization between consecutive calls to SHMEM collective routines. There are two special cases: * The shmem_barrier(3) routine allows the same pSync array to be used on consecutive calls as long as the active PE set does not change. * If the same collective routine is called multiple times with the same active set, the calls may alternate between two pSync arrays. The SHMEM routines guarantee that a first call is completely finished by all PEs by the time processing of a third call begins on any PE. Because the SHMEM routines restore pSync to its original contents, multiple calls that use the same pSync array do not require that pSync be reinitialized after the first call. SHMEM Function Inlining Some SHMEM functions that can be called from C/C++ are defined in the form of macros in the mpp/shmem.h header file. These functions are inlined by default on some platforms. To deactivate the automatic inlining of SHMEM functions from C/C++, add the following option to your C/C++ command line: -D_SHMEM_MACRO_OPT=0. SHMEM Application Placement on NUMA Systems On non-uniform memory access (NUMA) systems, such as Origin series systems, SHMEM start-up processing ensures that the process associated with a SHMEM PE executes on a processor near the memory associated with a SHMEM PE. The following environment variables allow you to control the placement of the SHMEM application on the system: Variable Description PAGESIZE_DATA Specifies the desired page size in kilobytes for program data areas. Specify an integer value. On Origin series systems, supported values include 16, 64, 256, 1024, and 4096. SMA_BAR_COUNTER Specifies the use of a simple counter barrier algorithm. By default, this variable is not enabled for jobs with PE counts of 64 or more. SMA_BAR_DISSEM Specifies the use of the alternate barrier algorithm, the dissemination/butterfly, within the shmem_barrier_all(3) function. This alternate algorithm provides better performance on jobs with larger PE counts. The SMA_BAR_DISSEM option is enabled for jobs with PE counts of 64 or higher. By default, this variable is not enabled for jobs with PE counts below 64. SMA_DBX Specifies the PE number to be debugged. If you set SMA_DBX to n, PE n prints a message during program startup, describing how to attach to it with the DBX debugger. PE n sleeps for seven seconds. If you set SMA_DBX to n,s, PE n will sleep for s seconds. SMA_DPLACE_INTEROP_OFF Disables a SHMEM/dplace interoperability feature available beginning with IRIX 6.5.13. By setting this variable, you can obtain the behavior of SHMEM with dplace on older releases of IRIX. By default, this variable is not enabled. SMA_DSM_CPULIST Specifies a list of CPUs on which to run a SHMEM application. To ensure that processes are linked to CPUs, this variable should be used in conjunction with SMA_DSM_MUSTRUN. For an explanation of the syntax for this environment variable, see the section entitled "Using a CPU List." SMA_DSM_MUSTRUN Enforces memory locality for SHMEM processes. Use of this feature ensures that each SHMEM process will get a CPU and physical memory on the node to which it was originally assigned. This variable has been observed to improve program performance on IRIX systems running release 6.5.7 and earlier, when running a program on a quiet system. With later IRIX releases, under certain circumstances, setting this variable is not necessary. Internally, this feature directs the library to use the process_cpulink(3) function instead of process_mldlink(3) to control memory placement. SMA_DSM_MUSTRUN should not be used when the job is submitted to miser (see miser_submit(1)) because program hangs may result. By default, this variable is not enabled. The process_cpulink(3) function is inherited across process fork(2) or sproc(2). For this reason, when using mixed SHMEM/OpenMP applications, it is recommended either that this variable not be set, or that _DSM_MUSTRUN also be set (see p_environ(5)). SMA_DSM_OFF When set to any value, deactivates processor-memory affinity control. When set, SHMEM processes run on any available processor, whether or not it is near the memory associated with that process. SMA_DSM_PPM When set to an integer value, specifies the number of processors to be mapped to every memory. The default is 2 on Origin 2000 systems. The default is 4 on Origin 3000 systems. SMA_DSM_TOPOLOGY Specifies the shape of the set of hardware nodes on which the PE memories are allocated. Set this variable to one of the following values: Value Action cube A group of memory nodes that form a perfect hypercube. NPES/SMA_DSM_PPM must be a power of 2. If a perfect hypercube is unavailable, a less restrictive placement will be used. cube_fixed A group of memory nodes that form a perfect hypercube. NPES/SMA_DSM_PPM must be a power of 2. If a perfect hypercube is unavailable, the placement will fail, disabling NUMA placement. cpucluster Any group of memory nodes. The operating system attempts to place the group numbers close to one another, taking into account nodes with disabled processors. (Default for IRIX 6.5.11 and higher). free Any group of memory nodes. The operating system attempts to place the group numbers close to one another. (Default for IRIX 6.5.10 and earlier releases). SMA_DSM_VERBOSE When set to any value, writes information about process and memory placement to stderr. SMA_INFO Prints information about environment variables that can control libsma execution. SMA_SYMMETRIC_SIZE Specifies the size, in bytes, of symmetric memory. This is the size of static space plus per-PE symmetric heap size. SMA_VERSION Prints the libsma library release version. Using a CPU List You can manually select CPUs to use for a SHMEM application by setting the SMA_DSM_CPULIST shell variable. This is treated as a comma and/or hyphen delineated ordered list, specifying a mapping of SHMEM processes to CPUs. The shepherd process is not included in this list. Examples: Value CPU Assignment 8,16,32 Place three SHMEM processes on CPUs 8, 16, and 32. 32,16,8 Place the SHMEM process rank zero on CPU 32, one on 16, and two on CPU 8. 8-15,32-39 Place the SHMEM processes 0 through 7 on CPUs 8 to 15. Place the SHMEM processes 8 through 15 on CPUs 32 to 39. 39-32,8-15 Place the SHMEM processes 0 through 7 on CPUs 39 to 32. Place the SHMEM processes 8 through 15 on CPUs 8 to 15. Note that the process rank is the value returned by _my_pe(3I). CPUs are associated with the cpunum values given in the hardware graph(hwgraph(4)). The number of processors specified must equal the number of SHMEM processes (excluding the shepherd process) that will be used. If an error occurs in processing the CPU list, the default placement policy is used. Using dplace(1) The environment variables described previously allow you to map SHMEM processes and memories with hardware processors and nodes. The dplace(1) command, which is available on Origin series systems, can give you additional control over application placement. Perform the following steps to use the dplace(1) command with SHMEM programs: * Create file placefile with these contents: threads $NPES + 1 memories ($NPES +1)/2 in topology cube distribute threads 1:$NPES across memories * Execute your program with NPES set to the number of PEs. For example, to run with 4 PEs, invoke your program this way: env NPES=4 dplace -place placefile a.out Interoperability SHMEM routines can be used in conjunction with MPI message passing routines in the same application. Programs that use both MPI and SHMEM should call MPI_Init and MPI_Finalize but omit the call to the start_pes routine. SHMEM PE numbers are equal to the MPI rank within the MPI_COMM_WORLD environment variable. On IRIX clustered systems, you can use SHMEM to comunicate only with processes running on the same host. Use the shmem_pe_accessible function to determine if a remote PE is accessible via SHMEM communication from the local PE. Compiling SHMEM Programs The SHMEM routines reside in libsma.so. The following sample command lines compile programs that include SHMEM routines: * IRIX systems: cc -64 c_program.c -lsma CC -64 cplusplus_program.c -lsma f90 -64 -LANG:recursive=on fortran_program.f -lsma f77 -64 -LANG:recursive=on fortran_program.f -lsma * IRIX systems with Fortran 90 version 7.2.1 available: f90 -64 -LANG:recursive=on -auto_use shmem_interface fortran_program.f -lsma The shmem_interface module is intended for use only with the -auto_use option. This module provides compile-time checking of interfaces. The keyword=arg actual argument format is not supported for SHMEM subroutines defined in the shmem_interface procedure interface module. The IRIX N32 ABI, selected by the -n32 compiler option, is also supported by SHMEM, but is recommended only for small process counts and program memory sizes, due to the limitation in the size of virtual addresses imposed by the N32 ABI. The use of the N64 ABI, selected by the -64 compiler option, is recommended for most SHMEM programs. Program Start-up The SHMEM implementation uses mapped files to render static memory remotely accessible on IRIX systems. The result is that enough file space must be available in /var/tmp to accommodate a file of size npes * staticsz, where npes is the number of PEs and staticsz is the size of the program's static data area. Static data includes Fortran common blocks and C/C++ static data. If a SHMEM program's memory requirements exceed available file space in /var/tmp, a SHMEM run-time error message is generated. You can use the TMPDIR environment variable to select a directory in a file system with sufficient file space. To minimize SHMEM program start-up time, use symmetric memory allocated by the SHPALLOC(3F) or shmalloc(3C) routines instead of static memory. Memory allocated by these routines does not require a corresponding file space allocation in /var/tmp. This avoids problems when file space is low and executes more quickly when start-up processing needs to handle large static memory areas. NOTES The SHMEM software is packaged with the Message Passing Toolkit. ENVIRONMENT VARIABLES For information on environment variables that affect SHMEM routines, see the "SHMEM Application Placement on NUMA Systems" section of this man page. EXAMPLES Example 1. The following Fortran SHMEM program runs on IRIX systems: PROGRAM REDUCTION REAL VALUES, SUM COMMON /C/ VALUES REAL WORK CALL START_PES(0) VALUES = MY_PE() CALL SHMEM_BARRIER_ALL ! Synchronize all PEs SUM = 0.0 DO I = 0,NUM_PES()-1 CALL SHMEM_REAL_GET(WORK, VALUES, 1, I) ! Get next value SUM = SUM + WORK ! Sum it ENDDO PRINT*,'PE ',MY_PE(),' COMPUTED SUM=',SUM CALL SHMEM_BARRIER_ALL END Since start_pes(3) is called with a value of 0, the number of PEs used to run the program is specified by the NPES environment variable. This Fortran program directs all PEs to sum simultaneously the numbers in the VALUES variable across all PEs. By executing the program using the following command line, you can you can run the program with 4 PEs: env NPES=4 a.out Example 2. The following C SHMEM program runs on IRIX systems: #include <mpp/shmem.h> main() { long source[10] = { 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 }; static long target[10]; start_pes(0); if (_my_pe() == 0) { /* put 10 words into target on PE 1 */ shmem_long_put(target, source, 10, 1); } shmem_barrier_all(); /* sync sender and receiver */ if (_my_pe() == 1) shmem_udcflush(); /* needed on T90 */ printf("target[0] on PE %d is %d\n", _my_pe(), target[0]); } In this C program, PE 0 sends 10 integers to the target array on PE 1. By executing the program using the following command line, you can you can run the program with 2 PEs: env NPES=2 a.out SEE ALSO dplace(1) The following man pages also contain information on SHMEM routines. See the specific man pages for implementation information. CC(1), cld(1), f90(1), f90(1M), mpprun(1) shmem_add(3), shmem_and(3), shmem_barrier(3), shmem_barrier_all(3), shmem_broadcast(3), shmem_cache(3), shmem_collect(3), shmem_cswap(3), shmem_event(3), shmem_fadd(3), shmem_fence(3), shmem_finc(3), shmem_get(3), shmem_iget(3), shmem_inc(3), shmem_iput(3), shmem_ixput(3), shmem_lock(3), shmem_max(3), shmem_min(3), shmem_mswap(3), shmem_my_pe(3), shmem_or(3), shmem_prod(3), shmem_put(3), shmem_quiet(3), shmem_short_g(3) shmem_short_p(3), shmem_stack(3), shmem_sum(3), shmem_swap(3), shmem_wait(3), shmem_xor(3), start_pes(3) shmalloc(3C) shpalloc(3F) MY_PE(3I), NUM_PES(3I) For information on using SHMEM routines with message passing routines, see the Message Passing Toolkit: PVM Programmer's Manual, or the Message Passing Toolkit: MPI Programmer's Manual.