PI2FD
Usage: PI2FD dest,src Modifies flags: None
PI2FD is a vector instruction that converts a vector register containing signed, 32-bit integers to single-precision, floating-point operands. When PI2FD converts an input operand with more significant digits than are available in the output, the output is truncated.
Packed Doubleword Integer to Single-Precision FP Convert
PI2FD mm1,mm2/m64 ; 0F 0F /r 0D [PENT,3DNOW]
PF2ID converts two signed 32-bit integers in the source operand to single-precision FP values, using truncation of significant digits, and stores them in the destination operand.
Example:
pi2fd mm1 mm2
PFSUBR
Usage: PFSUBR dest,src Modifies flags: None
Performs subtraction of the destination operand from the source operand. Both operands are single-precision, floating-point operands with 24-bit significands.
Packed Single-Precision FP Reverse Subtract
PFSUBR mm1,mm2/m64 ; 0F 0F /r AA [PENT,3DNOW]
PFSUBR subtracts the single-precision FP values in the destination from those in the source, and stores the result in the destination operand.
dst[0-31] := src[0-31] - dst[0-31],
dst[32-63] := src[32-63] - dst[32-63].
EXAMPLE:
pfsubr mm1 mm2
PFSUB
Usage: PFSUB dest,src Modifies flags: None
Performs subtraction of the source operand from the destination operand. Both operands are single-precision, floating-point operands with 24-bit significands.
Packed Single-Precision FP Subtract
PFSUB mm1,mm2/m64 ; 0F 0F /r 9A [PENT,3DNOW]
PFSUB subtracts the single-precision FP values in the source from those in the destination, and stores the result in the destination operand.
dst[0-31] := dst[0-31] - src[0-31],
dst[32-63] := dst[32-63] - src[32-63].
EXAMPLE:
pfsub mm1 mm2
PFRSQRT
Usage: PFRSQRT dest,src Modifies flags: None
Returns a low-precision estimate of the reciprocal square root of the source operand. The single result value is duplicated in both high and low halves of this instruction's 64-bit result.
Packed Single-Precision FP Reciprocal Square Root Approximation
PFRSQRT mm1,mm2/m64 ; 0F 0F /r 97 [PENT,3DNOW]
PFRSQRT performs a low precision estimate of the reciprocal square root of the low-order single-precision FP value in the source operand, storing the result in both halves of the destination register. The result is accurate to 15 bits. Negative operands are treated as positive operands for purposes of reciprocal square root computation, with the sign of the result the same as the sign of the source operand.
For higher precision reciprocals, this instruction should be followed by two more instructions: PFRSQIT1 and PFRCPIT2. This will result in a 24-bit accuracy. For more details, see the AMD 3DNow! technology manual.
EXAMPLE:
pfrsqrt mm1 mm2
PFRSQIT1
Usage: PFRSQRT1 dest,src Modifies flags: None
Performs the first intermediate step in the Newton-Raphson iteration to refine the reciprocal square root approximation produced by the PFSQRT instruction (the second and final step completes the iteration and is accurate to 24 bits).
Packed Single-Precision FP Reciprocal Square Root,
First Iteration Step
PFRSQIT1 mm1,mm2/m64 ; 0F 0F /r A7 [PENT,3DNOW]
PFRSQIT1 performs the first intermediate step in the calculation of the reciprocal square root of a single-precision FP value. The first source value mm1 is the square of the result of a PFRSQRT instruction, and the second source value mm2/m64 is the original value.
For the final step in a calculation, returning the full 24-bit accuracy of a single-precision FP value, see PFRCPIT2. For more details, see the AMD 3DNow! technology manual.
EXAMPLE:
pfrsqrt1 mm1 mm2
PFRCPIT2
Usage: PFRCPIT2 dest,src Modifies flags: None
Performs the second and final intermediate step in the Newton-Raphson iteration to refine the reciprocal or reciprocal square root approximation produced by the PFRCP and PFSQRT instructions, respectively.
Packed Single-Precision FP Reciprocal/ Reciprocal Square Root, Second Iteration Step
PFRCPIT2 mm1,mm2/m64 ; 0F 0F /r B6 [PENT,3DNOW]
PFRCPIT2 performs the second and final intermediate step in the calculation of a reciprocal or reciprocal square root, refining the values returned by the PFRCP and PFRSQRT instructions, respectively.
The first source value (mm1) is the output of either a PFRCPIT1 or a PFRSQIT1 instruction, and the second source is the output of either the PFRCP or the PFRSQRT instruction. For more details, see the AMD 3DNow! technology manual.
Example:
pfrcpit2 mm1 mm2
PFRCPIT1
Usage: PFRCPIT1 dest,src Modifies flags: None
Performs the first intermediate step in the Newton-Raphson iteration to refine the reciprocal approximation produced by the PFRCP instruction (the second and final step completes the iteration and is accurate to 24 bits).
Packed Single-Precision FP Reciprocal, First Iteration Step
PFRCPIT1 mm1,mm2/m64 ; 0F 0F /r A6 [PENT,3DNOW]
PFRCPIT1 performs the first intermediate step in the calculation of the reciprocal of a single-precision FP value. The first source value mm1 is the original value, and the second source value mm2/m64 is the result of a PFRCP instruction.
For the final step in a reciprocal, returning the full 24-bit accuracy of a single-precision FP value, see PFRCPIT2. For more details, see the AMD 3DNow! technology manual.
Example:
pfrcpit1 mm1 mm2
PFRCP
Usage: PFRCP dest,src Modifies flags: None
Returns a low-precision estimate of the reciprocal of the source operand. The single result value is duplicated in both high and low halves of this instruction's 64-bit result. The source operand is single-precision with a 24-bit significand.
Packed Single-Precision FP Reciprocal Approximation
PFRCP mm1,mm2/m64 ; 0F 0F /r 96 [PENT,3DNOW]
PFRCP performs a low precision estimate of the reciprocal of the low-order single-precision FP value in the source operand, storing the result in both halves of the destination register. The result is accurate to 14 bits.
For higher precision reciprocals, this instruction should be followed by two more instructions: PFRCPIT1 and PFRCPIT2. This will result in a 24-bit accuracy. For more details, see the AMD 3DNow! technology manual.
Example:
pfrcp mm1 mm2
PFPNACC
Usage: PFPNACC dest,src Modifies flags: None
Performs mixed negative and positive accumulation of the two doublewords of the 'dest' operand and the source operand. Then stores the results in the low and high words of the 'dest' operand, respectively.
Packed Single-Precision FP Mixed Accumulate
PFPNACC mm1,mm2/m64 ; 0F 0F /r 8E [PENT,3DNOW]
PFPNACC performs a positive accumulate of the two single-precision FP values in the source register and a negative accumulate of the destination register. The result of the accumulate from the destination register is stored in the low doubleword of the destination, and the result of the source accumulate is stored in the high doubleword of the destination register.
Both operands are single-precision, floating-point operands with 24-bit significands.
The operation is:
dst[0-31] := dst[0-31] - dst[32-63],
dst[32-63] := src[0-31] + src[32-63].
Example:
pfpnacc mm1 mm2
PFNACC
Usage: PFNACC dest,src Modifies flags: None
Performs negative accumulation of the two doublewords of the 'dest' operand and the 'src' operand. Then stores the results in the low and high words of the 'dest' operand.
Packed Single-Precision FP Negative Accumulate
PFNACC mm1,mm2/m64 ; 0F 0F /r 8A [PENT,3DNOW]
PFNACC performs a negative accumulate of the two single-precision FP values in the source and destination registers. The result of the accumulate from the destination register is stored in the low doubleword of the destination, and the result of the source accumulate is stored in the high doubleword of the destination register.
The operation is:
dst[0-31] := dst[0-31] - dst[32-63],
dst[32-63] := src[0-31] - src[32-63].
Both operands are single-precision, floating-point operands with 24-bit significands.
Example:
pfnacc mm1 Label