وو

وحید آنلاین . آرشیو وبلاگ وحیدمی دات آی آر . شرکت بیان. vahidmy.blog.ir

وو

وحید آنلاین . آرشیو وبلاگ وحیدمی دات آی آر . شرکت بیان. vahidmy.blog.ir

Using the Disassembler

       

Using the Disassembler  ...



When you [Open] a PE  without a Source Code inside or not written with RosAsm, RosAsm offers to Disassemble it. The proposed options are:


Normal Disassembly. This is the default, for the Source building, that is a simple Assembly Source. All Data and Code Labels are in the form of, for example, 'Code0403058', 'Data0405062'.


With Commented Hexa Code. In this Mode, the Hexa Code is given, in Comments, at the right of each Instruction.


With Symbolic Analyses. In this Mode, RosAsm tries to point out the Parameters passed on the Stack, for each Api call. When found, it replaces the mechanic labels by their true Names, as found in the Win32 Documentation. This is a first step toward full HLL interpretations.


Group the Data at Top. Many Executable Files have Data stored in the Code Section. With this Flag set 'On', RosAsm groups them all at the Top of the Source. With this Flag set 'Off', the Data will be given at the place where they are found, in the .Code Section.


Use a Map File. If RosAsm has already disassembled, say MyFile.exe, a MyFile.map has been saved, aside. Reusing it may speed up the Disassembling. When such a matching Map File is found aside the Disassembled Application, the default is set 'On'.



General approach


RosAsm's Disassembler is first, an Automatic Disassembler, that tries to provide a Source that could be re-compiled without any further hand work. This is actually effective on most small Demos. Between, say, 100 and 300 Ko, this may also work, but it depends, essentially on the quality (clean vs dirty construct) of the PE. Over this size (Megas) there is no hope, and  probably never will be, unless the PE organization would be absolutely standard.



What the Disassembler actually does


Intelligent Recognition of the PE's Sections, even in cases of merged Sections.


Recovering of all Resources (but Version Info Resources, not yet implemented in RosAsm). The Resources saved by Named IDs -instead of Numbered IDs are computed, but the RosAsm Resources Editors are not able, actually, to assume them (all RosAsm Resources Editors work only with Resources saved by Numbers). For the Main Menu, the original IDs are replaced by the usual RosAsm Equates Names, if the 'MainWindowProc' branchings are identified.


The various Data Formats recognitions, for Floats, Strings, pointers to Code or Data, are implemented.


Most small Applications, like Iczelion Tutorials Demos (all) and Test Department ones (all but Tut_5), Four-F ''Cocomac'' Demos, and so on... are correctly disassembled and re-assembled (re-run) in two Clicks. Even middle Size Applications, like the Iczelion Demo 35, for a RichEdit Editor, or Test Department's biggest Demos, seem to run fine, or, at least..., partially..., without any intermediate hand work between the Disassembling and the [Run].


A first HLL Interpretation, based on the Api calls Parameters may be applied. In this case, all the identified Api call parameters are replaced by the names found in the Api List Documentation. This process, actually based on the final Source text manipulations, is... very slow.


MainWindowProc and Main are detected and provided in the Source.


Api calls performed through a Jump Table (two Instructions instead of one) are replaced by the usual RosAsm direct calls. The original Api Jumps Table is provided, for cases of moves to Variables. In such cases, the Jumps Table Label is used.


A bit of Interactivity has been introduced with RosAsm V.2.022a. See Disassembler_Flags.



What it does not do


It will fail on encrypted PEs, on Auto-writeable Code, and on Code making a direct usage of hard Coded References, instead of Pointers.


It does not yet take care of the Menus-Items Equates for Dialogs. Only the first Menu, considered to be the default MainWindow-Menu is assumed.


Another weak point, is with the Recognition of small Chunks of Data nested inside Code. The Intelligent recognition may fail at deciphering if the Chunk is Data or not-called-Code. In such cases, it provides the DB Bytes, plus several commented Interpretations. (In other cases, when the Chunk is big enough to be identified true Data, the Chunk is moved into the normal Data).


The replacements of Structures Members Names and of Win32 Equates Names is not yet implemented.


The HLL Constructs (If, While, and friends) replacements are not yet implemented.



Practice


In practice, if you believe that you will have the possibility of disassembling a big Executable, and of re-Assembling it in two clicks, you will be disappointed. This is not at all the purpose of this Disassembler, and no Disassembler on earth will ever do that. It is simply impossible, unless the complete file would be 100% standard, from a Sections point of view and 100% clean, which is extremely uncommon.


So, work first with small Applications.


With middle size (100 / 300 Ko), you may have a valid Disassembly, that would not reflect exactly the Disassembled PE, because of minor failures.


The most usual failure cases are with erroneous  interpretations of small Chunks of Data or Code. In these cases, you may give a try to the [Bad Disassembly] Option of the Foat-Menu, when double-clicking on the suspected Label.


Then, once the Application is correctly re-compiled, it may also misbehave because of several minor points, that you may have to fix by hand, after analysis.



Purpose and scope


The Disassembler will remain under intensive development for several months. In its final state, it will be a Decompiler outputting a complete restoration of the Targeted File, ready for re-compilation in a significant amount of cases, which will make RosAsm a Programmer's Tool without any competitor, in that area.


Even in case of failure of the full ''Two-Clicks-Disassembler-ReAssembler'' process, the results will often be usable, at least, for study and for helping at the translation works.


The Disassembler is a Study and Translation Tool designed for the Open Source Movement. The main goal is to make the translations of Demos and Tuts to RosAsm syntax, as easy and fast as possible. Even when having the Sources, a port to Assembly may be not so easy, with big files. We can be sure that the Disassembler will, at least,  make fewer translation errors, and will take much less of our working time than we would when translating it by hand.


By no means do I intend to develop a Disassembler-Reassembler able to analyse any weird or on-design tricky PE, made with the purpose to resist  Disassembly.

 


~~~~~~~






Debugger

Debugger (Ludwig Hähne) .


Introduction


RosAsm comes with an integrated debugger which is built on top of the Win32 Debug API. When you Run / F5 your application (the debuggee) from inside RosAsm, it is automatically debugged. If you try to run a DLL, the debugger asks for a host process, which is expected to load the library.

 

The debugger will point out eventual exceptions in your source with the faulty instruction highlighted and a detailed exception description.

 

Furthermore you can set breakpoints in your source either at design time or at run time. When the debuggee encounters such a breakpoint, the OS transfers control to the debugger and halts all threads of the debuggee. 

With the debug dialog you can view the current flag states, register & data label values and view the contents of the whole address space of your application. The flags are embedded in the toolbar and can be shown and hidden through its context menu (right-click on toolbar).


Exceptions

When the exception box pops up, something went wrong in your application. The debug dialog title shows 'EXCEPTION' and the exception dialog title tells about the code module in which the crash occurred. A detailed exception description is given in the text window. Furthermore the instruction that caused the exception and its address are provided. In case of access violations additional information about the address which was tried to access and the access mode is shown below.


In general, the debuggee must be terminated when an exception occurs. Take care that the debug dialog is closed when you press Terminate. Therefore be sure to analyze the cause which may have led to the crash before you exit.


When you make use of structured exception handlers (SEH) the exception is still reported but you have the chance to forward it to the handler by Pass to SEH. If the exception is handled, the debuggee continues, otherwise the exception will be reported again without the possibility to pass it to the handler.


If the exception dialog caption doesn't show the name of your application but some other module like user32.dll, the exception happened outside of your application and (hopefully) a call to a external routine is highlighted. This does not mean, however, that it isn't your fault :) Most of the time missing or wrong parameters are the reason for these crashes. Check the call stack if in doubt which parameters have been passed to the routine(s).


Another possibility is an Exception in the Non-Code Section. The instruction pointer (EIP) was corrupted and triggered an access violation when the CPU tried to execute code at an inaccessible address. Instructions that may corrupt EIP are stray jumps or a ret when a wrong number of arguments have been passed. A look at the call stack might give a hint.


Registers


The register tab gives insight to the contents of the CPU registers. The debugger checks whether MMX, SSE is supported on your machine and shows additional pages in the tab if appropriate. Segment selectors and debug register (+EIP) pages  can viewed / hidden in the debug dialog menu settings. The combo-box offers various representations of the register contents, particularly useful to debug MMX/SSE code with a vector representation of mm0-7 / xmm0-7. 


The general purpose registers page differs from the other pages in that it contains buttons for each entry: If EBX contains a valid 32-bit virtual address in the process address space, clicking on the EBX-button takes you to the address referenced by EBX in the memory inspector.

 

The register contents are all shown zeroed until an exception occurred or a breakpoint is reached.


Setting Breakpoints

You can insert Breakpoints into your source, in two ways: 


Write int 3 (or int3) at the desired location. These are static breakpoints, represented by a 0CC Byte really inserted inside your PE Code Section, like any other Instruction. You cannot deactivate static breakpoints once the debuggee is running but you can switch off 'Hold on breakpoints' in the debug dialog menu settings to switch off all breakpoints.


Insert one or more dynamic breakpoint(s). These are breakpoints that are not represented inside your real Code. Instead, the Debugger inserts (and removes) them, on the fly, while your Application is being debugged. You can define such dynamic breakpoints by a simple mouse double-click, on the left margin of the source editor. When you double-click, a float menu offers you options for inserting/removing a breakpoint. If you use a small font, clicking exactly upon the very first left empty space (the margin), may be difficult. So, another option is available, for the same action: F4, that also runs the dynamic breakpoints float menu, and proposes the insertion at the beginning of the caret line. Note: When the debugger is running your application, you can add/delete dynamic breakpoints (whereas you cannot edit your source, at that time). 


Flow control / Tracing

When a breakpoint is encountered the next, not yet executed instruction is highlighted. To continue you can use the Continue menu items, the toolbar buttons or the corresponding shortcuts.


Run / F6 lets the debuggee continue to run without interruption through the debugger.

Step Into / F7 executes one instruction and then transfers control back to the debugger. Take care when stepping into API calls, some Windows versions (95 family) won't like or even allow it. When stepping through external module code you will see the module name in the debug dialog caption.

If the next instruction is a 'CALL' you have the possibility to Step Over the call, which allows the debuggee to run until the call has returned. This is also possible for looped instructions 'REPxx …'. F8 is even effective if the menu item is not available, having the same effect as Step Into, sometimes referred as 'auto-step-over'.

With Return to Caller / Ctrl+F7 you can step out of the current code and return to the caller in the process' code. This is useful if you're lost in deeply nested calls, or outside of the process' code. It won't work if the current code was called by the OS, like 'Main' or any 'WndProc'.

Terminate Debuggee / Ctrl+F6 lets you kill the debuggee at any time. First it kindly asks the debuggee to exit, if this does not happen within a few seconds the debuggee is terminated the hard way. It's also used to close the debuggee after an exception has occurred.


Source editor integration

RosAsm debugger operates on the source level. What does that mean in the context of assembly language? It means that you have full access to all symbols (code & data labels, equates) and the tracing takes place in the source editor. When stepping, the instruction which is executed next is highlighted. In case of instructions which have been generated by macros or pre-parsers the statement is highlighted from which the instruction was generated. To keep track of the progress inside the statement, the disassembled instruction is shown in the caption of the debug dialog. 


If single-stepping multiple instruction statements is not wished, you can switch to 'Source level stepping' in the debug dialog menu settings. In this mode the debuggee is continued until the next source statement is reached.

One of the benefits of operating on the source level, is the possibility of mouse sensitive data observation. When you move the mouse over an addressing expression (e.g. D$eax+8) in the source editor while debugging, you'll see the resolved virtual address (e.g. 010008 if eax=010000) and (if it is a valid address), the 32-bit value at this address in various data representations (hex, unsigned & signed decimal). Observable expressions start with D$, W$, B$, F$, R$, T$ and may contain registers, numbers, plain data labels & equates, segment selectors and '+', '-', '*' as operators. Q$, X$ and U$ are not yet supported. The size specifier determines the quantity and quality of the displayed memory contents. For example, D$ and F$ both reference 32-bit values but the latter is represented as floating point. Expressions which contain local labels or equates can only be observed if those belong to the procedure currently being executed. In other words, when you step through 'Foobar' you can observe the labels and equates local to Foobar in the form 'D$@Local' or the more common 'D@Local'. Examples for legal observable statements:

B$eax

F$DataLabel+ecx*4+EQUATE

W@Local+2 ; only if CurrentLabel@Local is defined

D$fs:8

To view the locals of the caller function(s) you can use the call stack described later in this document.


Data viewer

The data viewer shows all data symbols and their virtual addresses declared in your source. When selecting a symbol you can see the content with different representations in the window below the label list. The representations comprise Dword, Word, Byte sizes in Hexadecimal and Decimal (signed and unsigned) notation, floating point in single and double precision, and, if the data stream consists only of printable chars, the ASCII representation.

 

When right-clicking on a data symbol you can choose to view the content in the memory inspector, or, if the Dword content of the data is a valid address in the process' address space you can view the referenced memory. You can also search the declaration in the source, change the sort order of the symbols (by name, by address) or set watchpoints.


Watchpoints

Watchpoints can be assigned to data symbols. They are useful to observe write and/or read accesses to data, therefore they are sometimes referred as data breakpoints. To set a watchpoint, right click the symbol you want to observe in the data viewer and select 'Break On Write Access' or 'Break On Read/Write Access'. 


Watched data symbols are highlighted red (write) or orange (read/write). In the current implementation you cannot set multiple Watchpoints at the same time. Therefore, if you assign another watchpoint to a different symbol you will lose the old watchpoint.

 

When a watched access is observed the debuggee is halted, the title shows 'WP ...', the data viewer is activated and the watched data symbol is selected. Some implementation specific details:


Watchpoints utilize hardware debug register that are not handled correctly under old Windows version. Do not use watchpoints on Windows 95/98!

Access means, that at least one of the first 4 bytes (starting at the data address) is written to or read from.

Watchpoints only work on Dword aligned data.


Memory inspector

With the memory inspector you can view the memory contents of the allocated memory of your process. The memory is displayed in 4kB chunks which corresponds to the typical page size on x86 systems. The edit box shows the virtual address of the page in hex notation. Each list item contains an offset (e.g. +3F8) and the memory contents at this address. 

To view content at a specific address just enter it in the edit box and press return or use the virtual page table to select another region. You can also use segment overrides: e.g. FS:8 displays the TEB and goes to offset 8. 


The memory is shown aligned on 8 byte boundaries.


Call stack

The call stack shows the called procedures (labels) along with their parameters and local data. As the name implies it is derived from the stack content. When right-clicking on a function name you can show the invocation or declaration. The call stack is built using advanced interpretation mechanisms and should also show function calls inside modules, functions which don't setup stack-frames, ... However, it is only an interpretation. If your code or the modules you use make dynamic stack allocations (sub esp eax) or use jump tables the success rate will drop significantly. (It also does not handle spaghetti code very well) .


Best results are achieved if you follow these rules: 


Always use ret to return to the caller and to remove parameters from the stack 

Only use jumps to navigate inside your functions and not across your whole source 

Adhere to the standard code sequence to enter functions with local data (push ebp | mov ebp esp | sub esp x)


Function calls which belong to different modules (referenced code is outside your source) are grayed for clarity. If the information given is yet too detailed you can filter the output by right-clicking on any function name and selecting 'Hide module calls' or 'Hide intra-module calls'.


Debug log

The traditional way to debug code when no debugger is available is to log information to the console or a file. This might even make sense if using a debugger: For example, when the applications working is time-dependent and halting the program for inspection is not feasible because it would tamper with the output. Win32 offers a function for applications to pass strings to a possibly attached debugger: 'OutputDebugString'. When the debuggee calls 'OutputDebugString' the debugger is invoked and adds the string to the log tab and a log file is created aside the application with the name '[AppName]_dbg.log'. 

call 'Kernel32.OutputDebugString' {'Hello big brother' 0}

For convenience the log tab also lists mapped & unmapped modules and the creation and destruction of threads. Note that  'OutputDebugString' causes a context-switch to the debugger and thus is an expensive operation.


Address space

The address space tree shows all user accessible virtual memory pages of the debuggee. These are comprised of the mapped PE, the process environment block (PEB), the thread environment blocks (TEB), the stacks, the imported modules, the modules loaded by LoadLibrary, memory allocated by VirtualAlloc and the environment. 


The root nodes can be regarded as the 'logical groups' in which the memory was reserved while the leaf nodes represent the actual 4kB pages represented through the virtual start address and page properties (eXecute, Read, Write, Copy on write, Guard, No cache). When you double-click on a leaf node the page is loaded in the memory inspector.


~~~~~~~

EntryPoint


EntryPoint .



The default EntryPoint of any RosAsm Source is:


Main:


In case there would be some incompatibility (we have seen the occurency of a DLL that had to have to export a 'Main' Function...), you can redefine the EntryPoint with:


PREPARSE EntryPoint TheNewName


See Pre_Parser_concept, for PREPARSE rigid syntax, and take care that this EntryPoint implementation is not really a true PreParser, but just a dirty hack that does the required substitution.


As opposed to the other PREPARSE statements, this one cannot be mixed with the other ones. So, you must leave it stand alone on its own line. Example:


PREPARSE Equal Alternates

PREPARSE EntryPoint TheNewName


~~~~~~~







IncIncluder_Parser


IncIncluder_Parser .



To enable it, you have to write:


PREPARSE IncIncluder


See Pre_Parser_concept, for PREPARSE rigid syntax. In most other Assemblers, when modifying a .inc Source File, that is ''include''ed in several Applications Sources, you may update all of these Applications at once. The ''IncIncluder'' Pre-Parser's purpose is to provide such an implementation by modifying a TITLE in as many Applications, that may make use of it.



Syntax:


PREPARSE incIncluder


INCINCLUDE  D:\Path\Title.inc


Your Source must, of course, also have a:


 TITLE Title


... reflecting the ''Title.inc'' File.


Each time you will re-compile your Source, if the ''Title.inc'' File has been modified, the Pre-Parser will substitute the modified File to your actual Title Part.


The File Max Size is 1 Mb.


All Statements are Case Sensitive: 'INCINCLUDE' must be upper case and the File Name cases must fit with the TITLE Name Cases.


When saving a TITLE Part with [Ctrl]/[S], if the actual Part is the one targeted by an INCINCLUDE Statement, you will be proposed to either save to the original Name.inc, or to save, in the usual way, the Name.asm into the current Directory. This allows you to eventually directly modify an .inc File from any PE making use of it, after modification in the Source Editor.


This implementation, while offering the main advantage of a classical inc Method, fully preserves all of the so important advantages of Mono-File Programming. Notice that, as opposed to the classical inc Method, this implementation is not recursive. This is to say, that an inc File can only be a TITLE part, and cannot include other sub-inc Files.


In order to save you from the usual management problems coming with external Files, you should provide the full Path and Name for your .inc File, in your INCINCLUDE Statement. Example:


PREPARSE incIncluder


INCINCLUDE  D:\Programming\RosAsm\Includes\Macros.inc


And not use the default Directory with a lazy and dangerous:


INCINCLUDE  Macros.inc


Doing it this way will, also, much reduce the conflicts possibilities, if you distribute your Application Open Sources, as others could, as well, also have a different Macros.inc File in their default Directory, for example... In case of conflict, your Application could not be recompiled on the other Computer - it would be modified -, whereas, when not finding the required .inc File, the Parser simply does nothing at all, this is to say that the Application is directly re-compiled from the actual Source as it comes.


When distributing an Application making use of this feature, you do not need to include the inc File(s), as it already is in the actual Source.


~~~~~~~







Includer_Parser


Includer_Parser (Author: Kenny)   ..



To enable it, you have to write:


PREPARSE BinIncluder


See Pre_Parser_concept, for PREPARSE rigid syntax. The ''Includer'' Pre-Parser's purpose is to include  directly a File (any kind of File) directly into the .Data section of your PE.



Syntax.


PREPARSE BinIncluder


[BinInclude MyFile.ext:]


Once done, in your usual Source, you may access the Data stored at the 'MyFile.ext:' Label the usual way.


The usual 'Len' is also  automatically generated for providing the Length of the included Data. In this example case,  - as the Parser does it for you -, you do not have to write, at the end:


... MyFile.ext_Len: len]


The File Max Size is 1 Mb.


~~~~~~~







Equal_Parser


Equal_Parser  (Author: Scarmatil)  ..



To enable it, you have to write:


PREPARSE Equal


See Pre_Parser_concept, for PREPARSE rigid syntax. The ''Equal'' Pre-Parser's purpose is to imitate the HLL writing Style (non-Assembly syntax). ''Equal'' is, in fact, not one, but two different Pre-Parsers: The Statements Equal Pre-Parser, and the FPU Equal pre-Parser.



Compile-Time Expressions


Statements the Equal Pre-Parser translates, expressions like:


eax = 1

edx = &TRUE+(2*2)

D$Value1 = D$Value2

W$Value1+2 = W$Value2+2

D$Handle = call 'DLLNAME.DllFunction', Para1, ...  ; or any Proc call returning eax.


into:


mov eax 1

mov edx &TRUE+(2*2)

push D$Value2 | pop D$Value1

push W$Value2+2 | pop W$Value1+2

call 'DLLNAME.DllFunction', Para1, ...  | mov D$Handle eax



Run-Time Expressions


It also can parse Run-time Expressions (Expressions expanded into Instructions and, eventually Data Declarations, whose values will only be known at Run-Time):


eax = (D$a+D$b)/6+(D$c*7)

D$val3 = (3 * D$a + 12) / (24 * D$b) - D$c * D$c * 18 / 13 + D$z


R$FpuValue = 13.333+5*R$MyReal/((-1.24+F$MyFloat*W$MyWord)/3-(2+D$MyDWord))


Operators List:


=> ^        ; power operator, the first operand has to be strictly positive 

=> * / + - : common operators

   => // **    ; signed operators 


integrated FPU functions: cos(), sin(), abs(), tan() and sqrt().


ebx = sin(45)

T$FloatNumber = sqrt(abs(7+-3*ebx)+R$FpVal)


Signed operations:


To force signed integer division and multiplication you can use '//' and '**' operators. Notice that if one of the operands is a floating point value (6.2,T$MyVal,...) the operation will be signed even with '*' and '/' operators.


Conversions:


The parser can perform any conversion as long as they are not impossible (B$val1 = D$val2 ;for example):


  > ST0 = (D$a+D$b)/6+(D$c*7)       ; integer result stored in a FPU register

  > Q$MyQValue = B$MyByteVal  

  > ...


Array support:


There is no possible way to use arrays (X can be any of the data sizes stated below):

    - X$(address)     : this will be interpreted as : X$address*(X_size)

    - X$Array(index)  : this will be interpreted as : X$Array+index*(X_size)


[MyTable: D$ 'ZERO' 'ONE ' 'TWO ' 'THREE' 'FOUR' 'FIVE']


    mov eax  4

    eax  =  D$MyTable(eax)  ; eax = 'FOUR'



Supported data size:


(both in source and destination): B$, W$, D$, Q$, F$, R$ and T$.



 Basically this parser works as follows 


It first copies the operation from CodeSourceA to CodeSourceB ( from the first parenthesis to EOI)


Then spaces are added (if needed) before and after each parenthesis, operator and data  (registers and constants included)


The parsing begins and, the first closing parenthesis is searched. Each time the parser meets an opening parenthesis, its address is stored in ebx . Once a closing parenthesis is found, the operations in these parentheses are processed.  (for more details, see in each part)

  


Example of how the parser expands a Statement


   > eax = (D$b-D$a*(ebx+7/2)-3)


You can see what the Equal Parser does, with the Double-Click Float  Menu, Option [Unfold], just like with a User Defined Macro, if you Double-Click, say, on the above 'eax':

  

    FINIT

    PUSH EAX

    PUSH EBX

    PUSH EDX

    [AAAAAAAA: 0]

    MOV D$AAAAAAAA EBX

    [BAAAAAAA:T$ 3.5E+0000]

    FLD T$BAAAAAAA

    FIADD D$AAAAAAAA

    FSTP T$BAAAAAAA

    FLD T$BAAAAAAA

    FIMUL D$A

    FSTP T$CAAAAAAA

    [CAAAAAAA:T$ 0]

    FLD T$CAAAAAAA

    FISUBR D$B

    FSTP T$DAAAAAAA

    [DAAAAAAA:T$ 0]

    [EAAAAAAA:T$ 3]

    FLD T$EAAAAAAA

    FLD T$DAAAAAAA

    FSUBRP ST0 ST1

    FSTP T$EAAAAAAA

    POP EDX

    POP EBX

    POP EAX

    [FAAAAAAA: 0]

    FLD T$EAAAAAAA


~~~~~~~










Alternates_Syntaxes



Alternates_Syntaxes  ...


Beginners,   DO NOT READ THIS. The RosAsm syntax is clean, clear and simple. There is absolutely no reason for wishing to use another one. But older programmers are  inclined by human natures habit to do things in some ways and they hate to have to change, even from bad to good.


So, I have implemented a Pre-Parser that allows several alternate syntaxes. To enable them, you have to say:


PREPARSE Alternates


See Pre_Parser_concept for PREPARSE rigid syntax.



For sizes markers, instead of the clean and easy:


  mov  ecx  D$Counter


you can write, as well, any of the old and stupid forms of:


mov  ecx  D[Counter]

mov  ecx  dWord [Counter]

mov  ecx  dWord Ptr [Counter]


Same, of course, for all Sizes Markers:


D$Synbol    >    D / Dword /  Dword Ptr[Symbol]

W$Synbol    >   W / word    /  Word Ptr[Symbol]

B$Synbol     >    B /  Byte     /  Byte Ptr[Symbol]

Q$Synbol    >    Q / Qword /  Qword Ptr[Symbol]

R$Synbol     >    R /  Real    /   Real Ptr[Symbol]

F$Synbol     >    F /  Float    /  Float Ptr[Symbol]

T$Synbol     >    T / TByte   /  Tword / Tbyte Ptr / TWord Ptr [Symbol]

O$Synbol    >    O / Oword /  Oword Ptr[Symbol]

X$Synbol     >    X /  Xword /  Xword Ptr[Symbol]


(for Tbyte, Tword is an alternates historical stupidity).



Implicit size, when RosAsm can guess it from a Register Parameter is allowed too:


mov  ecx  [Counter]

mov  ax  [Counter]

mov  [Counter]  bl


Sizes can only be computed from Byte/Word/dWord Registers.



When declaring Data, you can as well replace the original D$, W$, B$, and so on, by DD, DW, DB, and so on, but this last formulation absolutely requires spaces separators (and nothing else) before and after the Marker (whereas the basic syntax does not).



When declaring Equates, if you can't live with:


[TRUE  1,    FALSE  0]


you can as well write (well,... this one is a good choice, IMHO):


[TRUE  =  1,    FALSE  =  0]


or even:


[TRUE  EQU  1,    FALSE  EQU  0]


These two formulations, too, require spaces separators before and after the '='  and  'EQU' dummy statement. This small limitation allows the Basic Syntax to reuse  these reserved words for  another purpose when they are standing just after/before ''[ / ]''. Example:


[=  =  e,    >  =  a, .........]




If you want to do something stupid such as Typings (the most confusing feature ever implemented in an Assembler), you can do it with Macros. For Example:


[Data | {#1Ptr: #2>L} | {#1  #2#1Ptr}]


Data  dWordData D$,  24  13  3


Note the comma after ''D$'', absolutely required because the Macro parser would confuse ''D$  24'' with ''D$24''. Note too, that, in such cases of Macros Declarations, as the Macros Parser does its job a long time after the Pre-Parser, you can not have usage of ''DD/DB/DW'' alternates.


mov  eax  dWordData             ; eax = 24 !!!!!!!!!!!!

mov  ebx  dWordDataPtr        ; ebx = Adress of  dWordData!!!!!!!!!!!!!!



Or even more stupid:


[Data|{ADDR#1: #2>L} | {#1  #2ADDR#1}]


mov  eax  Mydata+8          ; eax = 3 !!!!!!!!!!!!

mov  ebx  ADDR_Mydata ; ebx = Adress of  dWordData!!!!!!!!!!!!!!


Your source will just be completely unreadable, the Right-Click search will not find the symbols declarations,  your development time will be twice as long and the compilation will be twice as long also. But it works...

~~~~~~~


Pre-Parsers concept


Pre-Parsers concept  ...



RosAsm being a 'True Assembler', means that it does not include any internal hidden Macros that would enhance the basics of Assembly Syntax, and that all HLL features are to be implemented via user defined Macros. You are not constrained in any way, other than the microprocessor's instruction set, and, are in full control at all times.


When you are writing a set of .If / .Else / .EndIf, in, say, MASM, you are using a Compiler, because these features are hidden inside, being an integral part of the Compiler itself. When you are doing exactly the same thing with RosAsm, you are using a ''True Assembler'', because you defined the Macros in order to enable this.


But, RosAsm being a true Assembler does not mean that it may not include selectable Pre-Parsers, that the user may -or not- run in order to include the HLL statements he may want, in his Sources, just the reverse of what most HLLs do with Inline Assembly. Even if, one day, a volunteer wants to write a Basic Pre-Parser for RosAsm, you will be allowed to run it with:


> PEPARSE Basic


And, this will not change an inch about RosAsm being and remaining a True Assembler, in its internals, and running as such by default. So  'PREPARSE' tells what Pre-Parser(s) to run when compiling your Source, adding alternates or HLL capabilities on the Top of the Assembler, just like any external ''Front-End'' would do. Though it might seem a bit of a strange formulation, RosAsm is internally open to Pre-Parsers.


Four Pre-Parsers are actually available. See Alternate_Syntaxes , Equal_Parser, and BinIncluder_Parser, IncIncluder_Parser.


I hope that my way of introducing these Pre-Parsers makes it as clear as possible. Once you run a Pre-Parser you are no longer running a ''True Assembler'', but, at least, a Compiler, if not a real HLL, and the way it is done perfectly respects the Bottom-Up Approach and organization of RosAsm.


Like the TITLE reserved word (see 'TITLE' in Source_Editor), the PREPARSE Statement is to be written in upper case only, must begin at the first Position on a Line, and cannot hold anything else than what is expected.


PREPARSE writing convention is:


PREPARSE Option(s)     


; where 'Option(s) ' may actually be: 'Alternates' , 'Equal' and 'Includer':


PREPARSE Alternates

PREPARSE Equal

PREPARSE Includer

PREPARSE Alternates, Equal, Includer


The purpose of this recently added KeyWord is to allow the implementations of any number of Pre-Parsers, without any Compilation time penalty, for the Sources making no use of these Parsers. In the future, we will possibly add an OOA Pre-Parser, a Basic Pre-Parser, and so on, without any quantity limitation and with possible conflicts management in case of multi-selections (for example, if we implement various conflicting flavors of HLL Pre-Parsers).


~~~~~~~





XMM


XMM   ..



Most XMM instructions apply on Packed Data which are 128 bits long. This is to say, in practice, packets of 2 qWords, or 4 dWords, or 8 Words, or 16 Bytes.


For Memory pointing to this kind of Data, the standard keyword is 'OctoWord'. What is implemented, in RosAsm as, for example:


addpd  XMM0  O§MyOctoWord


As 'O' may easily be confused with 'Q', and as the pointed memory is not really an OctoWord but a Packed of Data (none being OctoWord sized), in accordance with the NASM developers, I have implemented too 'X' (for Xmm, for undefined,...):


addpd  XMM0  X§MyOctoWord


These two notations are available. 'X' notation is also available for other ''irregular'' sizes addressing (FPU).


       FSAVE  >>>  108 bytes mem

       FLDENV >>>    28 bytes mem

       FRSTOR >>>  180 bytes mem

       FSTENV >>>    28 bytes mem



The comparison instructions for XMM (CMPPS / CMPPD / CMPSS / CMPSD) have a 3 members encodage, with a really demential encodage:


cmpps  XMM1  XMM2  imm8


''imm8'' is a immediate value (from 0 to 7) telling what kind of comparison is to be performed, which are:


1 = EQ (Equal)

2 = LT (Less Than)

3 = LE (Less Than or Equal) UNORDed

4 = NEQ (Not Equal)

5 = NLT (Not Less Than)

6 = NLE (Not Less or Equal)

7 = ORD (Ordered)


The Mnemonics forms I implemented for these are, (example for: CMPPS, same for the others):


CMPPSEA / CMPPSLT / CMPPSLE

CMPPSNEQ / CMPPSNLT / CMPPSNLE

CMPPSORD /CMPPSUNORD


... More readable with '_':


CMP_PS_EA / CMP_PS_LT / CMP_PS_LE

CMP_PS_NEQ / CMP_PS_NLT / CMP_PS_NLE

CMP_PS_ORD / CMP_PS_UNORD


And, for compatibility with other Compilers:


CMP_EA_PS / CMP_LT_PS / CMP_LE_PS_LE

CMP_NEQ_PS / CMP_NLT_PS / CMP_NLE_PS

CMP_ORD_PS / CMP_UNORD_PS


Which achieves into _3_ (!!!...) different forms for these awfully designed Instructions, that are all computed by the Assembler. Example of 3 forms producing the same encodage.


CMPPD xmm1 X$Data 5

CMP_PD_NLE xmm1 X$Data

CMP_NLE_PD xmm1 X$Data



Declaring Data with 'X$' / 'O$' doesn't make any sense as Octowords Data size do not really exist. Declare the Data for what they really are. XMM Packed Data must be 128 bits aligned:


[<16 MyPackedData: D$ 0 0 0 0]

[<010 MyPackedData: D$ 0 0 0 0]

[<0010000 MyPackedData: D$ 0 0 0 0]

~~~~~~~



XMM

MMX instructions


MMX instructions  ....



MMX registers are noted:


        MM0, MM1, ..., MM7



MMX registers and memory operands are always 64 bits long (8 bytes) but can operate on 1 Qword, 2 Dwords, 4 Words or 8 bytes.


When you find in the documentation:       PADD  (Add with wrap-around)    for example, this instruction has actually 3 forms, one for bytes, one for  words, and one for dWords; so that this mnemonic is extended to 3 forms in RosAsm:


      PADD        PADDB / PADDW / PADDDD


As usual, you can write more clearly:


P_AAD_B / ...   , that could stand for 'Packed Addition on Bytes / ...


Each time an MMX instruction has several forms, if you do not specify any in the mnemonic, RosAsm will consider it to be the smallest possible one (P_ADD = P_ADD_B). If you do specify one for a mnemonic which has only one form (no need of any specifier), you won't get an error message, as far as your parameter specifier is fitting.


~~~~~~~