Equates .
RosAsm offers four variants and forms of Equates:
User defined Equates.
OS Equates.
Text Equates.
Body Equates.
User defined Equates
[true 1, false 0, MayBe 00101010101010101]
Are user defined equates. Equates cannot be redefined; double declarations are not allowed. They are simple substitutions, for any instance of evocations, in the Statements. You can use equates evocations inside macros declarations, data, and code. Even though it is not really useful, you can as well use Equates Evocations inside Equates Declaration:
[aaa 012, bbb aaa+1, ccc bbb+2]
mov eax bbb ; eax = 13
sub eax ccc ; eax = -1
OS Equates
mov eax &TRUE
The leading '&', of &TRUE indicates that this is an OS Equate. You do not have to declare anything, because RosAsm includes a complete encoded Table of OS Equates.
How it works:
In your RosAsmFiles Folder, there is a File called Equates.equ, that actually contains more than 60,000 Equates Declarations. The first time you run RosAsm, it searches for this file, and, if found, it computes it, and outputs two more Files: Equates.nam and Equates.num. The first one is a Table of dWords CheckSums of the Names. The second one is a Table of dWords for the values of all Equates Names. The Assembler holds these two Tables in Memory when compiling your Application.
When you use such an OS Equate, The Assembler, first, recomputes the CheckSum of the Name and searches for this Checksum, in the Names-CheckSums Table. Then, the value is found at the corresponding Displacement, in the Values Table.
The purpose of this particular implementation is, first, to save you from any boring include, - or from having to declare these Equates, in your Source -, second to speed-up the Compilation.
The User Defined Equates and the OS Equates are two independent things. For example, even though there is an OS Equate saying that &TRUE is 1, you can, as well, declare a User Equate saying that [TRUE 3], and there will be no conflict.
When downloading an update of the Equates.equ File, after having saved it, in your RosAsmFiles Directory, you have to update your Equates.nam and Equates.num Files by the Menu [Tool] / [Rebuild Equates] Option.
If you wish to implement some Equates of your own as OS Equates, this is possible: You simply save some MyEquates.equ File aside Equates.equ, you run the [Rebuild Equates] Option, and this is done. Though, this possibility is really not recommended, as therefore, nobody but you will be able to recompile your Appplications. Modifying the list of OS Equates should be reserved to the Maintainers of RosAsm only.
Usage
For 'multi-Equates' parameters, the internal operation is always OR unless you specify some '+imm' after the evocation. You can write as well:
&MB_ICONINFORMATION + &MB_SYSTEMMODAL ; readable
&MB_ICONINFORMATION+&MB_SYSTEMMODAL ; less readable
&MB_ICONINFORMATION___&MB_SYSTEMMODAL ; _____hhhmmmm_____
&MB_ICONINFORMATION&MB_SYSTEMMODAL ; unreadable but shorter
But not:
&MB_ICONINFORMATION__&MBSYSTEMMODAL ; > Missing '_' in 'MB_SYSTEMMODAL'
The Win Equates Parser considers '+' '_' or 'nothing', between 2 &Equates as an OR operation and not as an addition, because this is a frequent error to list 2 different Equates that have same values (and meaning) but with different misleading names.
Note that this does not apply to immediates. example:
mov eax &TRUE+1 ; >>> eax = 2 (and not 1)
You can easily verify any Win Equates values by Right-Clicks on the symbol (it will fail if not in upper case because the same fast search routine is used here, as in encodage). Heavy use of Win32 Equates does not spoil any compilation time. This replacement is blazingly fast.
Text Equates
Text Equates are nothing but the Text form of User Defined Equates:
[SomeText 'Some text declared by Equate' ]
[SomeData: SomeText , 0] ; is same as:
[SomeData: 'Some text declared by Equate' , 0 ]
Even multi-Lines 'Text' is allowed (with double quotes, of course).
The single or double quotes are part of the Equate replacement.
Body Equates
Body Equates are also User Defined Equates. They can contain anything, in between angle Brackets:
[BodyEquate < 'Some Text' , 0, 0654, &NULL > ]
The delimiters of Body Equates are: < > . This implies, of course, that you cannot make any use of these two Characters inside the Equate body, even inside declared 'Text'. In the upper example, the two spaces (after <, and before >), are not part of the Equate body and are stripped off by the Parser. The other ones are. This is to prevent from logical problems, in cases of evocations requiring Spaces: The leading space is to be given in the Evocation.
This feature is useful, for example, to declare Expressions inside Equates:
[SomeEquate < (24 shl 2) > ]
The two differences between Text Equates and Body Equates are that Text Equates include the Quotes, while Body Equates do not include the angle Brackets, and that Body Equates do not include the leading and ending spaces , whereas, Text Equates do.
~~~~~~~
Compile-Time Functionalities .
Equates
Macros_Keys
Macros_Basics
Inside_Macros
Internal_Strings
Internal_Counters
Conditional_Macros
Macros_Examples
ParaMacros
The Compile-Time Engines are the set of implementations that enables the programmer with defining his own source substitutions. In other words, how to redefine the Assembly language defaults.
As well as for Data, these Declarations are to be given inside Square Brackets.
Equates
The Equates engine is the simpler one: It just replaces the defined words by something else, typically on a one-to-one basis:
[TRUE 1]
This Equate declaration means: 'Replace all intances of 'TRUE' by '1'.
Macros
The Macros engine also offers a replacement functionality, like the Equates Engine, but, whereas the Equates can be applied to any member of an Instruction, the Macros apply only on the first element of a Statement, and, whereas, the Equates are 'one word only' Macros can have Parameters, and can produce complex multi-Statements outputs. Also, a Macro can be applied to the first component of a Statement Typically:
[push | push #1 | #+1]
push eax, ebx, ecx
will be interpreted as:
push eax | push ebx | push ecx
Macros are usually very rarely used by Asm programmers because, in other Assemblers, they are difficult to learn and to use . As RosAsm tends to be as low level as possible, I have taken particular care to develop powerful, easy to learn and easy to use Macros. So, RosAsm can be considered both as low level in its internals and as a high level Assembler by user defined Macros.
Not considering the error check (which has been really a never ending work in RosAsm development), this Macros system is the feature in which I have spent the most of my working time. To get some idea of the power of this system, just take a look at the ''EndP'' Macro in Beginners' Tut5 and what it does.
Paramacros
The ParaMacros Engine does about the same things as the Macros Engine, but applies to a Parameter of a Statement:
[RGB | (#1 or (#2 shl 8) or (#3 shl 16))]
mov eax {RGB 011, (011*2), {ReverseByte 0}}
Same as:
mov eax 0_FF_22_11
Note:
The Declared Equates, Macros and ParaMacros apply to the whole Source, whatever positions of the Declarations and of the Evocations. You may, as well, have the Declarations at the very end of a Source making use of these substitutions functionalities.
~~~~~~~
Structures ..
C-Like Structures
Some bad Assemblers have a 'STRUCT' KeyWord performing hidden Declarations and Substitutions. This feature is particularly bad and incompatible with assembly. It has only inconvenience and no advantages. When one writes, with these Assemblers, something like:
POINT STRUCT
x DWORD ?
y DWORD ?
POINT ENDS
; ...
LOCAL Ps :POINT
; ...
mov eax, Ps.x
MyPoint <POINT>
mov ebx MyPoint.y
What is effectively done by the Assembler would be expressed, in RosAsm syntax, by something like this:
[X = 0 Y = 4]
Local @Point 8
mov eax D@Point+X
[MyPoint: D$ ? #2]
mov ebx D$MyPoint+Y
This is to say that this 'STRUCT' KeyWord performs a hidden Declaration of 'Displacement Equates', and that the '.x' formulation is unfolded as an added Displacement from the given Base.
This feature is very bad because the real symbols appear nowhere in your Source. No Search or Debug feature can point out, in your Sources, Declarations that do not exist. For Win32 Structures, compared to the way RosAsm gives you all of this by the [Struct] Dialog, there is no advantage. For your own private Structures, there is no advantage either as you have, in all cases, to declare these Structures by yourself, by hand.
As long as C-Like Structures Declarations are nothing but hidden Equates sets Declarations, there is no reason, in Assembler for not declaring them for what they are, instead of writing risky, unreadable and difficult to debug Code. Several users, with a previous experience with bad Assemblers, have asked me for having this 'Missing Feature' available. My answer has always been 'no' and will always be 'no', even at the cost of having users leaving and giving up on RosAsm. ( This has already happened several times ).
The [Struct] Dialog
The [Struct] menu item opens a Dialog for helping at Win32 Structures programming. You will retreive, in the ClipBoard, the selected Structure. The four top-Right Radio-Buttons define the way the items names are formatted.
The three downward left CheckButtons provide 3 different forms for the Structures: The common one (Data), another one written in Equates format (useful to point to Tables given in return by Win Notify messages-), and a last one, useful when you wish to declare a Structure (and the pointers to the data you wish to access) on the Stack (Example:
Structure @PAINTSTRUCT 64 , @hdc 0, ......
In that last case, avoid keeping all the symbols you don't use, to save RosAsm compiling task running for nop on Equates job.
And do never forget that, when writing such stack reservations by hand (for example because of a missing Structure in the integrated List), you *have* to align the Structures whole size on dWord Boundary. The Stack MUST remain aligned. Several Win Structures sizes are not aligned. The Structure given to you by the [Struc] Dialog are always properly size aligned.
The [Data to Equates] Dialog
One another Dialog is available in the [Tools] PopUp Menu. As, for you, writing the Structures under Data form is easier than calculating the relative displacements of each Member in a Structure (this may be a real pain with big and complex Sets), the purpose of this Dialog is to make this translation easy. The Left Edit Control receives you Data version input, and the right Edit Control gives you the Structure Equates translation.
When using this feature, it is a good idea to keep the Data version in your Source (in a Multi-Lines Comment -;;-), in order to not re-enter all of your Data each time you have to modify them.
Missing Structures
The available Structures in Structures.str Data File are only the basic ones of Win32. This simple set of Data is, in fact, a demential amount of work, done by a lot of programmers. So, I first thought of writing a set of tools to try to directly translate *.inc C Files into Asm syntax. Despite my efforts, I have been unable to implement this work for several reasons:
The first one is, of course that building these tools is a huge work, due to the weird and irregular syntaxes of all C files that might be found here and there. The fact that nobody ever did these translations tools, is not only because of the lack of volunteers. The method to try to understand what is a C Structure Declaration is yet to be found...
The second one is that, really coming over with it, would require learning C. I first wanted to do it, but as, on one hand, I hate C over all, and as, on the other hand, it is more and more evident to me that I cannot do everything by myself, I leave it open for others to implement.
Until ReactOS is available for public use, we will not have enough volunteers to do the job.
Until ReactOS is available for public use, we will not know with which C syntax we have to deal with, and as with one single C, the translation tools are almost impossible to write, with various C's it becomes absolutely impossible.
Note: I write 'tools', and not 'tool' because, after many weeks wasted at this, it now seems to me impossible to translate in one single pass, without any hand work between the various jobs' routines.
Waiting for better times, this is to say for the oncoming of a Asm volunteer programmer who will perfectly understand C and who will want to write the translation tools, if you are in need of alien Structures, the 'simpler' (!!!) way... is to do it by hand. As the amount of work to be done is widely out of human scale, if you perform such translations, think of forwarding them back to me (even if very partial) so that I can prepare future public releases of these Data under Asm form Files.
A partial translation for DirectX is available inside the Dx Demo available at my page.
~~~~~~~
Expressions ...
The Expressions Parser works with parenthesis, on immediate values, from inside to outside for nested levels.
The keys are: + - * / or and xor not shl shr. For integers.
Immediate Numbers can be given in Decimal, Hexa or Binary.
Examples:
mov eax (((654+12) and (00_1010_1010*2)) +3) ; eax = 19
mov eax (not 2) ; eax = 0_FFFF_FFFD
There is no precedence. All operations are performed, from inside to outside levels and inside one given level, from left to right.
The Expressions Parser is for immediates only. You can use Equates inside.
The Expressions Parser can hold Foating Point Unit Reals, but in a much more restrictive way than integers. The syntax is, for example:
[RealNumber: R§ (34/25.1 -2)]
Only the 4 Operations (- + * / ) in the case of Reals. You cannot use nested Parenthesis sets, for example, R§ ((34/25.1)-2)], or use logical Operators, will give rise to error Messages.
Win Equates and User Equates are allowed inside Expressions.
Expressions are allowed inside Equates.
~~~~~~~
Double Negation ...
RosAsm can do: 'double negative conditionals'. What's that ?
Let us suppose you want some 'DO / Loop_Until' Macro. In order to handle this negative case (we loop until the given condition is NOT true), you would usually have to organize your macro like this:
[Do | D0:] [Loop_Until | cmp #1 #3 | j#2 D1> | jmp D0<< | D1:]
Do
;
; some instructions
;
Loop_Until eax ne 0
>>>> D0:
; some instructions
cmp eax 0 | jne D1> ; double jumps just like
jmp D0<< ; in Local_Labels Select_Case example
D1:
With double negation of condition, you can now write the more accurate:
[Do | D0:] [Loop_Until | cmp #1 #3 | jn#2 D0<<]
Do
;
; some instructions
;
Loop_Until eax ne 055
>>>> D0:
; some instructions
cmp eax 055 | jnne D0<< ; (jnne = je) <<<<<< This is double negation <<<<<<
(Double negation is not a feature of the macros parser. It is a simple addition at beginning of the 'tttn Bits' conditions parser).
So, the upper Select_Case example given in Local_Labels section now becomes clean and still handles the negative cases:
;...
[Case | Jmp C9>> | C2: | cmp D§Select_Val #2 | jn#1 C2>]
;...
And the clean unfolding:
; ...
C2: cmp D§Select_Val 0 | jne C2>
; ...
; some code lines
; ...
jmp C9>>
C2: cmp D§Select_Val 1 | jne C2>
; ...
~~~~~~~
Data alignment ....
When the Processor accesses Memory, if the targeted Data are not aligned on their own boundary, it then has to perform 2 Memory accesses instead of one.
For beginners, aligned on their own boundary, means that a Word begins on a paired Address. Example:
At Address 0600301 >>> Word/dWord are not aligned.
At Address 0600302 >>> Word is aligned but dWord is not aligned.
At Address 0600304 >>> Word/dWord are aligned.
Each time you open a new square bracket for Data, RosAsm, by default, does the alignment job for you, so that, if you take care of writing your Data in decreasing size order inside one given Bracket set, all your Data are always properly aligned. To prevent RosAsm from doing this alignment, see Data_Management.
As a general rule, I am not much of a fanatic of optimizations, but this one is so simple and so important that it seems to me useful to implement it directly as a Data Encoder default behaviour.
~~~~~~~
Code Alignment ....
I first didn't think of implementing any Align feature for code as this seemed to me an 'out of purpose' idea. But some other programmers told me that it was an absolute must have in an assembler. So, here it is:
align 16 ; (decimal only parameter)
>>> inserts the needed number of NOPs to reach the desired boundary.
If the NOPs number is greater than 4, the two first NOPs are replaced by a short JMP.
Intel documentation gives only two cases for using Alignment:
A loop entry label should be 16-byte aligned when it is less than 8 bytes away from that boundary.
A label that follows an unconditional branch or function call should be 16-bytes aligned when it is less than 8 bytes away from that boundary.
I implemented this because of the aforementioned advice, but I am not yet fully convinced it is a good idea. About optimization, I am of the opinion that:
Optimization for one processor is not always true for another one.
Optimization for yesterdays Intel processors is not the one for today and much more so is not the one for tomorrow.
Optimization, today, is far too sophisticated to be considered in the real world, even if it works fine on some scholastic examples and if some champions insist on using it every day.
Processors are now blazing fast and clean assembler writing is enough.
If not, rewrite the code.
Optimizing assembler code while, at the same time, other programmers write entire systems or sound routines in C, is indeed a laughing matter.
The better code optimization is 'the code we do not write' ; I mean that a clean short and intelligent routine is always much faster than a dirty, stupid long one.
~~~~~~~
RosAsm Tables ..
When you want to declare several data at once, what other assemblers call DUP is done with '#n' (n can be given either in Decimal, in HexaDecimal or in Win Equate form):
[ValuesSet: 1 2 3 4 5 #10 ] ; same as:
[ValuesSet: 1 2 3 4 5 1 2 3 4 5 1 2 3 4 5 1 2 3 4 5 1 2 3 4 5
1 2 3 4 5 1 2 3 4 5 1 2 3 4 5 1 2 3 4 5 1 2 3 4 5]
[Hum: B$ 'Hum!!!...' 0 #3 ] ; same as:
[Hum: B$ 'Hum!!!...' 0 Hum: B$ 'Hum!!!...' 0 Hum: B$ 'Hum!!!...' 0]
Notice that the Looping Marker is after the Data Declaration. Not before.
'#n' data loop cannot be nested actually (I'll see that later). As this feature is mostly used for static tables declarations you will certainly use it usually in:
[My_Static_Table: 0 #128]
Avoid mixing Tables and normal Data. This might mislead you both at writing what you think and interpreting what you wrote. If you want to implement a 'data loop' inside a continuous set of data, do it with as many sets as needed with 'unaligned data sets' (see explanations in Data_Management)
[OneString: B$ 'One text', 0]
[<OneStaticTable: 0 #32]
[<AnotherString: B$ 'Another text', 0]
~~~~~~~
RosAsm Text ..
Either ' or '' as delimiter (as you wish, but not mixed):
[My_Text: b$ 'One day, One hand, one foot', 0] ; right
[My_Text: b$ 'One's day, One's hand, one's foot', 0] ; wrong
[My_Text: b$ ''One's day, One's hand, one's foot'', 0] ; right
[My_Text: b$ 'She told me ''One day, One hand, one foot'' ', 0] ; right
Differences between single and double quotes
Single quotes text cannot include CR/LF, double quotes can. The consequence of this is that you must consider single quotes as the standard feature (perfect control of open text errors) and double quotes as an alternate feature (bad control -RosAsm might not find error in some cases-) useful only for solving problems like upper examples.
Double quotes are useful too if you mean to include CR/LF in text declaration without having to write some 'B$ 13 10'. Once again, be aware that, with double quotes, RosAsm will be unable to detect open text errors in some cases. I will not modify this because flexibility and ease of use is much more important in my opinion than security.
Text expressions in code
mov eax 'abcd'
does what you mean (al = 'a' / ah = 'b' / ...). No need to reverse reading order.
Avoid using 'text' for values direct computations:
mov eax 'abcd'-'1234' ; Doesn't work (maybe later...)
mov eax 'abcd'+'1234' ; should work...
Tabs
In RosAsm edition, Tabs are now fully reserved for Indentation. If you want a true Tab inside Text Data, insert ASCII 9 bytes:
[TabbedText: B§ 'abcde', 9, 'abcde', 0]
~~~~~~~
RosAsm Numbers ...
24 ; this is 24 decimal (no leading zero)
018 ; this is 24 hexadecimal (one leading zero)
0011000 ; this is 24 binary (two leading zeros)
equivalent but best:
24
0_18
00_0001_1000
See real number at FPU
As RosAsm is intended is to be a true assembler it has of course none of MASM style extended data, like:
char for byte
integer for dword
boolean for byte
float for real4
IntPtr for far ptr integer
If a character is a byte, there is absolutely no reason to call it a 'char'.
Char and Boolean are _qualities_ / 'B$' points a _quantity_.
There is one exception for that: I implemented 'R$/...' data declarations for the FPU format to simplify both the parser job and user writings.
Some good programmers are in the opposite opinion and think that 'fashioned typings' increase readability. So, if you agree with them, simply set the equates you need for this:
[char B§]
mov al, char esi ; Stupid, but it works too...
~~~~~~~