Sunday, March 2, 2008

SAX Parser tips

Recently I got couple of interesting questions from my friends who are working on XML and using SAX parser to 'parse' the XML data - for performance and memory efficient; SAX parser can work efficiently even for 2 GB XML files!

Identifying Self ending tags:
Actually in XML both <br/> and <br/></br> are equivalent. So, using SAX parser you can't find whether it is a self ending tag or not. However there is a work around for it - using locator objects!

For <br/>, in both startElement and endElement you get the same location (getLineNumber() and getColumn number()) will be same.

For <br/></br>, they will be different – column numbers will be different (or even line number!).

But, using Locator object with SAXParser might slightly decrease the performance.
Also one more thing, all SAX may not support Locators as this is an optional feature.

More about Locators can be found at http://www.saxproject.org/apidoc/org/xml/sax/Locator.html


Handling default attributes

Problem:
Input file : <xhtml:td>VI</xhtml:td>Benzyl</xhtml:td>

Output file :
<xhtml:td rowspan="1" colspan="1">VI</xhtml:td>
<xhtml:td align="left" rowspan="1" colspan="1">Benzyl</xhtml:td>

The data has "rowspan" , “colspan” automatically included in the output. But the same is not present in the input.

The dtd declaration for the xhtml:td is as below
<!ATTLIST %td.qname;
%attrs;
abbr %Text; #IMPLIED
axis CDATA #IMPLIED
headers IDREFS #IMPLIED
scope %Scope; #IMPLIED
xhtml:rowspan %Number; "1"
xhtml:colspan %Number; "1"
%cellhalign;
%cellvalign;
>

These attributes are coming because they have a default value in DTD.

In the DTD it is mentioned that the default value of the xhtml:rowspan is 1, so unless you specify some value the rowspan will be 1.

Even if you don’t declare that attribute, SAXParser automatically get the value from the DTD (a ‘special’ feature of SAX parser called DTD defaulting).

You can only handle this in SAX2 parser (not in SAX parser version 1.x). I think most of the SAX parser available (like one comes with JDK1.5) today are SAX2.

In your startElement method, you will get an object of Attributes2 instead of Attributes; Actually Attributes2 is a subclass of Attributes.

Attributes2 interface has method isSpecified() which returns true unless the attribute value was provided by DTD defaulting.

So, keep this check in startElement method:



public void startElement (String uri, String localName,
String qName, Attributes attributes) throws SAXException
{
if (attributes instanceof Attributes2) {
Attributes2 att = (Attributes2) attributes
for (int i = 0; i < att.getLength(); i++) {
if (att.isSpecified(i)) // present in xml file
System.out.println(att.getQName(i) + "=\"" + att.getValue(i) + "\"");
else {// not present in xml file, came from DTD.
}
}
} // if not, we don't have a choice output all attributes.
}



There is another better way to check whether the SAX Parser Attributes2 or not - by checking the system property http://xml.org/sax/features/use-attributes2
More details at http://www.saxproject.org/apidoc/org/xml/sax/package-summary.html#package_description

Sunday, February 17, 2008

Compare two word documents using MS Word

Compare two word documents using MS Word
1. Open a document.
2. On the Tools menu, click Compare and Merge Documents.
3. Select the document that you want to compare to the copy that is currently open.
4. Click the arrow next to Merge, and then do one of the following:
* To display the results of the comparison in the selected document, click Merge.
* To display the results in the document that is currently open, click Merge into current document. * To display the results in a new document, click Merge into new document.

The differences will be displayed as comments in the new document.





For example here the Last updated date has been changed from 13-04-07 to
11-05-07.

Speed up the start of Acrobat Reader

Opening PDF files are taking time, then

Method 1:
Every time you run Adobe Acrobat, up to 20 plugins are loaded
unnecessarily - most users do not need even a fraction of them!
To disable unneeded plugins and make them optional instead, follow these
instructions:
1. Browse to the plugins folder: C:\Program Files\Adobe\Acrobat
7.0\Reader\plug_ins
2. Create a new folder named Optional
3. Move all files from the plug_ins folder to Optional, except
EWH32.api, print*.api, and Search*.api

Method 2:
Download the software Adobe Reader SpeedUp software from
http://www.tnk-bootblock.co.uk/software/, which does the same thing
specified in Method 1, but in a nice GUI friendly option.

Method 3:
Install a light weight PDF reader (other than Acrobat). The best free
PDF reader is Foxit Reader. No need to install, just copy the executable
and run from your system. This is very light and extremely fast.

Foxit reader is available at
http://us01.foxitsoftware.com/foxitreader/foxitreader.zip (just 1.8 MB
size).


Copyright (c) 2008 - Suresh