HTML parser, doesn't support HTML5.
More...
deprecated content model
Lookup the HTML tag in the ElementTable.
Lookup the given entity in EntitiesTable.
Lookup the given entity in EntitiesTable.
The HTML DTD allows a tag to implicitly close other tags.
The HTML DTD allows a tag to implicitly close other tags.
This is kept for compatibility with previous code versions.
Allocate and initialize a new HTML parser context.
Allocate and initialize a new HTML SAX parser context.
Create a parser context for an HTML in-memory document.
Parse an HTML document and invoke the SAX handlers.
Parse an HTML in-memory document.
Parse an HTML in-memory document and build a tree.
Create a parser context to read from a file.
parse an HTML file and build a tree.
Parse an HTML file and build a tree.
int
htmlUTF8ToHtml (unsigned char *out, int *outlen, const unsigned char *in, int *inlen)
Take a block of UTF-8 chars in and try to convert it to an ASCII plus HTML entities block of chars out.
int
htmlEncodeEntities (unsigned char *out, int *outlen, const unsigned char *in, int *inlen, int quoteChar)
Take a block of UTF-8 chars in and try to convert it to an ASCII plus HTML entities block of chars out.
Check if an attribute is of content type Script.
Set and return the previous value for handling HTML omitted tags.
Create a parser context for using the HTML parser in push mode.
Parse a chunk of memory in push parser mode.
Free all the memory used by a parser context.
Reset a parser context.
Applies the options to the parser context.
Applies the options to the parser context.
Convenience function to parse an HTML document from a zero-terminated string.
Convenience function to parse an HTML file from the filesystem, the network or a global user-defined resource loader.
xmlDoc *
htmlReadMemory (const char *buffer, int size, const char *URL, const char *encoding, int options)
Convenience function to parse an HTML document from memory.
xmlDoc *
htmlReadFd (int fd, const char *URL, const char *encoding, int options)
Convenience function to parse an HTML document from a file descriptor.
Convenience function to parse an HTML document from I/O functions and context.
Parse an HTML document and return the resulting document tree.
Parse an HTML in-memory document and build a tree.
Parse an HTML file from the filesystem, the network or a user-defined resource loader.
Parse an HTML in-memory document and build a tree.
Parse an HTML from a file descriptor and build a tree.
Parse an HTML document from I/O functions and source and build a tree.
HTML parser, doesn't support HTML5.
This module orginally implemented an HTML parser based on the (underspecified) HTML 4.0 spec. As of 2.14, the tokenizer conforms to HTML5. Tree construction still follows a custom, unspecified algorithm with many differences to HTML5.
The parser defaults to ISO-8859-1, the default encoding of HTTP/1.0.
- Copyright
- See Copyright for the status of this software.
- Author
- Daniel Veillard
◆ htmlParserOption
This is the set of HTML parser options that can be passed to htmlReadDoc, htmlCtxtSetOptions and other functions.
| Enumerator |
|---|
| HTML_PARSE_RECOVER | No effect as of 2.14.0.
|
| HTML_PARSE_NODEFDTD | Do not default to a doctype if none was found.
|
| HTML_PARSE_NOERROR | Disable error and warning reports to the error handlers.
Errors are still accessible with xmlCtxtGetLastError().
|
| HTML_PARSE_NOWARNING | Disable warning reports.
|
| HTML_PARSE_PEDANTIC | No effect.
|
| HTML_PARSE_NOBLANKS | Remove some text nodes containing only whitespace from the result document.
Which nodes are removed depends on a conservative heuristic. The reindenting feature of the serialization code relies on this option to be set when parsing. Use of this option is DISCOURAGED.
|
| HTML_PARSE_NONET | No effect.
|
| HTML_PARSE_NOIMPLIED | Do not add implied html, head or body elements.
|
| HTML_PARSE_COMPACT | Store small strings directly in the node struct to save memory.
|
| HTML_PARSE_HUGE | Relax some internal limits.
See XML_PARSE_HUGE in xmlParserOption.
- Since
- 2.14.0
Use XML_PARSE_HUGE with older versions.
|
| HTML_PARSE_IGNORE_ENC | Ignore the encoding in the HTML declaration.
This option is mostly unneeded these days. The only effect is to enforce ISO-8859-1 decoding of ASCII-like data.
|
| HTML_PARSE_BIG_LINES | Enable reporting of line numbers larger than 65535.
- Since
- 2.14.0
Use XML_PARSE_BIG_LINES with older versions.
|
| HTML_PARSE_HTML5 | Make the tokenizer emit a SAX callback for each token.
This results in unbalanced invocations of startElement and endElement.
For now, this is only usable to tokenize HTML5 with custom SAX callbacks. A tree builder isn't implemented yet.
- Since
- 2.14.0
|
◆ htmlAttrAllowed()
htmlStatus htmlAttrAllowed
(
const htmlElemDesc *
elt,
int legacy )
- Deprecated
- Don't use.
- Parameters
-
elt HTML element
attr HTML attribute
legacy whether to allow deprecated attributes
- Returns
- HTML_VALID
◆ htmlAutoCloseTag()
int htmlAutoCloseTag
(
xmlDoc *
doc,
The HTML DTD allows a tag to implicitly close other tags.
The list is kept in htmlStartClose array. This function checks if the element or one of it's children would autoclose the given tag.
- Deprecated
- Internal function, don't use.
- Parameters
-
doc the HTML document
name The tag name
elem the HTML element
- Returns
- 1 if autoclose, 0 otherwise
◆ htmlCreateFileParserCtxt()
Create a parser context to read from a file.
- Deprecated
- Use htmlNewParserCtxt and htmlCtxtReadFile.
A non-NULL encoding overrides encoding declarations in the document.
Automatic support for ZLIB/Compress compressed document is provided by default if found at compile-time.
- Parameters
-
filename the filename
encoding optional encoding
- Returns
- the new parser context or NULL if a memory allocation failed.
◆ htmlCreateMemoryParserCtxt()
Create a parser context for an HTML in-memory document.
The input buffer must not contain any terminating null bytes.
- Deprecated
- Use htmlNewParserCtxt and htmlCtxtReadMemory.
- Parameters
-
buffer a pointer to a char array
size the size of the array
- Returns
- the new parser context or NULL
◆ htmlCreatePushParserCtxt()
void * user_data,
const char * chunk,
int size,
const char * filename,
Create a parser context for using the HTML parser in push mode.
- Parameters
-
sax a SAX handler (optional)
user_data The user data returned on SAX callbacks (optional)
chunk a pointer to an array of chars (optional)
size number of chars in the array
filename only used for error reporting (optional)
enc encoding (deprecated, pass XML_CHAR_ENCODING_NONE)
- Returns
- the new parser context or NULL if a memory allocation failed.
◆ htmlCtxtParseDocument()
Parse an HTML document and return the resulting document tree.
- Since
- 2.13.0
- Parameters
-
ctxt an HTML parser context
input parser input
- Returns
- the resulting document tree or NULL
◆ htmlCtxtReadDoc()
const char * URL,
const char * encoding,
int options )
Parse an HTML in-memory document and build a tree.
See htmlCtxtUseOptions for details.
- Parameters
-
ctxt an HTML parser context
str a pointer to a zero terminated string
URL only used for error reporting (optional)
encoding the document encoding (optional)
- Returns
- the resulting document tree
◆ htmlCtxtReadFd()
int fd,
const char * URL,
const char * encoding,
int options )
Parse an HTML from a file descriptor and build a tree.
See htmlCtxtUseOptions for details.
NOTE that the file descriptor will not be closed when the context is freed or reset.
- Parameters
-
ctxt an HTML parser context
fd an open file descriptor
URL only used for error reporting (optional)
encoding the document encoding (optinal)
- Returns
- the resulting document tree
◆ htmlCtxtReadFile()
const char * filename,
const char * encoding,
int options )
Parse an HTML file from the filesystem, the network or a user-defined resource loader.
See htmlCtxtUseOptions for details.
- Parameters
-
ctxt an HTML parser context
filename a file or URL
encoding the document encoding (optional)
- Returns
- the resulting document tree
◆ htmlCtxtReadIO()
void * ioctx,
const char * URL,
const char * encoding,
int options )
Parse an HTML document from I/O functions and source and build a tree.
See htmlCtxtUseOptions for details.
- Parameters
-
ctxt an HTML parser context
ioread an I/O read function
ioclose an I/O close function
ioctx an I/O handler
URL the base URL to use for the document
encoding the document encoding, or NULL
- Returns
- the resulting document tree
◆ htmlCtxtReadMemory()
const char * buffer,
int size,
const char * URL,
const char * encoding,
int options )
Parse an HTML in-memory document and build a tree.
The input buffer must not contain any terminating null bytes.
See htmlCtxtUseOptions for details.
- Parameters
-
ctxt an HTML parser context
buffer a pointer to a char array
size the size of the array
URL only used for error reporting (optional)
encoding the document encoding (optinal)
- Returns
- the resulting document tree
◆ htmlCtxtReset()
Reset a parser context.
Same as xmlCtxtReset.
- Parameters
-
ctxt an HTML parser context
◆ htmlCtxtSetOptions()
Applies the options to the parser context.
Unset options are cleared.
- Since
- 2.14.0
With older versions, you can use htmlCtxtUseOptions.
- Parameters
-
ctxt an HTML parser context
- Returns
- 0 in case of success, the set of unknown or unimplemented options in case of error.
◆ htmlCtxtUseOptions()
Applies the options to the parser context.
The following options are never cleared and can only be enabled:
- Deprecated
- Use htmlCtxtSetOptions.
- HTML_PARSE_NODEFDTD
- HTML_PARSE_NOERROR
- HTML_PARSE_NOWARNING
- HTML_PARSE_NOIMPLIED
- HTML_PARSE_COMPACT
- HTML_PARSE_HUGE
- HTML_PARSE_IGNORE_ENC
- HTML_PARSE_BIG_LINES
- Parameters
-
ctxt an HTML parser context
- Returns
- 0 in case of success, the set of unknown or unimplemented options in case of error.
◆ htmlElementAllowedHere()
int htmlElementAllowedHere
(
const htmlElemDesc * parent,
- Deprecated
- Don't use.
- Parameters
-
parent HTML parent element
elt HTML element
- Returns
- 1
◆ htmlElementStatusHere()
htmlStatus htmlElementStatusHere
(
const htmlElemDesc *
parent,
const htmlElemDesc * elt )
- Deprecated
- Don't use.
- Parameters
-
parent HTML parent element
elt HTML element
- Returns
- HTML_VALID
◆ htmlEncodeEntities()
int htmlEncodeEntities
(
unsigned char * out,
int * outlen,
const unsigned char * in,
int * inlen,
int quoteChar )
Take a block of UTF-8 chars in and try to convert it to an ASCII plus HTML entities block of chars out.
- Deprecated
- Only supports HTML 4.
- Parameters
-
out a pointer to an array of bytes to store the result
outlen the length of out
in a pointer to an array of UTF-8 chars
inlen the length of in
quoteChar the quote character to escape (' or ") or zero.
- Returns
- 0 if success, -2 if the transcoding fails, or -1 otherwise The value of inlen after return is the number of octets consumed as the return value is positive, else unpredictable. The value of outlen after return is the number of octets consumed.
◆ htmlEntityLookup()
const htmlEntityDesc * htmlEntityLookup
(
const
xmlChar *
name )
Lookup the given entity in EntitiesTable.
- Deprecated
- Only supports HTML 4.
TODO: the linear scan is really ugly, an hash table is really needed.
- Parameters
-
name the entity name
- Returns
- the associated htmlEntityDesc if found, NULL otherwise.
◆ htmlEntityValueLookup()
const htmlEntityDesc * htmlEntityValueLookup
(
unsigned int value )
Lookup the given entity in EntitiesTable.
- Deprecated
- Only supports HTML 4.
TODO: the linear scan is really ugly, an hash table is really needed.
- Parameters
-
value the entity's unicode value
- Returns
- the associated htmlEntityDesc if found, NULL otherwise.
◆ htmlFreeParserCtxt()
Free all the memory used by a parser context.
However the parsed document in ctxt->myDoc is not freed.
- Parameters
-
ctxt an HTML parser context
◆ htmlHandleOmittedElem()
int htmlHandleOmittedElem
(
int val )
Set and return the previous value for handling HTML omitted tags.
- Deprecated
- Use HTML_PARSE_NOIMPLIED
- Parameters
-
val int 0 or 1
- Returns
- the last value for 0 for no handling, 1 for auto insertion.
◆ htmlInitAutoClose()
void htmlInitAutoClose
(
void )
◆ htmlIsAutoClosed()
int htmlIsAutoClosed
(
xmlDoc *
doc,
The HTML DTD allows a tag to implicitly close other tags.
The list is kept in htmlStartClose array. This function checks if a tag is autoclosed by one of it's child
- Deprecated
- Internal function, don't use.
- Parameters
-
doc the HTML document
elem the HTML element
- Returns
- 1 if autoclosed, 0 otherwise
◆ htmlIsScriptAttribute()
int htmlIsScriptAttribute
(
const
xmlChar *
name )
Check if an attribute is of content type Script.
- Deprecated
- Only supports HTML 4.
- Parameters
-
name an attribute name
- Returns
- 1 is the attribute is a script 0 otherwise
◆ htmlNewParserCtxt()
◆ htmlNewSAXParserCtxt()
Allocate and initialize a new HTML SAX parser context.
If userData is NULL, the parser context will be passed as user data.
- Since
- 2.11.0
If you want support older versions, it's best to invoke htmlNewParserCtxt and set ctxt->sax with struct assignment.
Also see htmlNewParserCtxt.
- Parameters
-
sax SAX handler
userData user data
- Returns
- the htmlParserCtxt or NULL in case of allocation error
◆ htmlNodeStatus()
- Deprecated
- Don't use.
- Parameters
-
legacy whether to allow deprecated elements (YES is faster here for Element nodes)
- Returns
- HTML_VALID
◆ htmlParseCharRef()
- Deprecated
- Internal function, don't use.
- Parameters
-
ctxt an HTML parser context
- Returns
- 0
◆ htmlParseChunk()
const char * chunk,
int size,
int terminate )
Parse a chunk of memory in push parser mode.
Assumes that the parser context was initialized with htmlCreatePushParserCtxt.
The last chunk, which will often be empty, must be marked with the terminate flag. With the default SAX callbacks, the resulting document will be available in ctxt->myDoc. This pointer will not be freed by the library.
If the document isn't well-formed, ctxt->myDoc is set to NULL.
Since 2.14.0, xmlCtxtGetDocument can be used to retrieve the result document.
- Parameters
-
ctxt an HTML parser context
chunk chunk of memory
size size of chunk in bytes
terminate last chunk indicator
- Returns
- an xmlParserErrors code (0 on success).
◆ htmlParseDoc()
Parse an HTML in-memory document and build a tree.
- Deprecated
- Use htmlReadDoc.
This function uses deprecated global parser options.
- Parameters
-
cur a pointer to an array of
xmlChar
encoding the encoding (optional)
- Returns
- the resulting document tree
◆ htmlParseDocument()
Parse an HTML document and invoke the SAX handlers.
This is useful if you're only interested in custom SAX callbacks. If you want a document tree, use htmlCtxtParseDocument.
- Parameters
-
ctxt an HTML parser context
- Returns
- 0, -1 in case of error.
◆ htmlParseElement()
This is kept for compatibility with previous code versions.
- Deprecated
- Internal function, don't use.
- Parameters
-
ctxt an HTML parser context
◆ htmlParseEntityRef()
- Deprecated
- Internal function, don't use.
- Parameters
-
ctxt an HTML parser context
str location to store the entity name
- Returns
- NULL.
◆ htmlParseFile()
xmlDoc * htmlParseFile
(
const char *
filename,
const char * encoding )
Parse an HTML file and build a tree.
- Parameters
-
filename the filename
encoding encoding (optional)
- Returns
- the resulting document tree
◆ htmlReadDoc()
const char * url,
const char * encoding,
int options )
Convenience function to parse an HTML document from a zero-terminated string.
See htmlCtxtReadDoc for details.
- Parameters
-
str a pointer to a zero terminated string
url only used for error reporting (optoinal)
encoding the document encoding (optional)
- Returns
- the resulting document tree.
◆ htmlReadFd()
const char * url,
const char * encoding,
int options )
Convenience function to parse an HTML document from a file descriptor.
NOTE that the file descriptor will not be closed when the context is freed or reset.
See htmlCtxtReadFd for details.
- Parameters
-
fd an open file descriptor
url only used for error reporting (optional)
encoding the document encoding, or NULL
- Returns
- the resulting document tree
◆ htmlReadFile()
xmlDoc * htmlReadFile
(
const char *
filename,
const char * encoding,
int options )
Convenience function to parse an HTML file from the filesystem, the network or a global user-defined resource loader.
See htmlCtxtReadFile for details.
- Parameters
-
filename a file or URL
encoding the document encoding (optional)
- Returns
- the resulting document tree.
◆ htmlReadIO()
void * ioctx,
const char * url,
const char * encoding,
int options )
Convenience function to parse an HTML document from I/O functions and context.
See htmlCtxtReadIO for details.
- Parameters
-
ioread an I/O read function
ioclose an I/O close function (optional)
ioctx an I/O handler
url only used for error reporting (optional)
encoding the document encoding (optional)
- Returns
- the resulting document tree
◆ htmlReadMemory()
xmlDoc * htmlReadMemory
(
const char *
buffer,
int size,
const char * url,
const char * encoding,
int options )
Convenience function to parse an HTML document from memory.
The input buffer must not contain any terminating null bytes.
See htmlCtxtReadMemory for details.
- Parameters
-
buffer a pointer to a char array
size the size of the array
url only used for error reporting (optional)
encoding the document encoding, or NULL
- Returns
- the resulting document tree
◆ htmlSAXParseDoc()
const char * encoding,
void * userData )
Parse an HTML in-memory document.
If sax is not NULL, use the SAX callbacks to handle parse events. If sax is NULL, fallback to the default DOM behavior and return a tree.
- Deprecated
- Use htmlNewSAXParserCtxt and htmlCtxtReadDoc.
- Parameters
-
cur a pointer to an array of
xmlChar
encoding a free form C string describing the HTML document encoding, or NULL
sax the SAX handler block
userData if using SAX, this pointer will be provided on callbacks.
- Returns
- the resulting document tree unless SAX is NULL or the document is not well formed.
◆ htmlSAXParseFile()
xmlDoc * htmlSAXParseFile
(
const char *
filename,
const char * encoding,
void * userData )
parse an HTML file and build a tree.
Automatic support for ZLIB/Compress compressed document is provided by default if found at compile-time. It use the given SAX function block to handle the parsing callback. If sax is NULL, fallback to the default DOM tree building routines.
- Deprecated
- Use htmlNewSAXParserCtxt and htmlCtxtReadFile.
- Parameters
-
filename the filename
encoding encoding (optional)
sax the SAX handler block
userData if using SAX, this pointer will be provided on callbacks.
- Returns
- the resulting document tree unless SAX is NULL or the document is not well formed.
◆ htmlTagLookup()
const htmlElemDesc * htmlTagLookup
(
const
xmlChar *
tag )
Lookup the HTML tag in the ElementTable.
- Deprecated
- Only supports HTML 4.
- Parameters
-
tag The tag name in lowercase
- Returns
- the related htmlElemDesc or NULL if not found.
◆ htmlUTF8ToHtml()
int htmlUTF8ToHtml
(
unsigned char * out,
int * outlen,
const unsigned char * in,
int * inlen )
Take a block of UTF-8 chars in and try to convert it to an ASCII plus HTML entities block of chars out.
- Deprecated
- Internal function, don't use.
- Parameters
-
out a pointer to an array of bytes to store the result
outlen the length of out
in a pointer to an array of UTF-8 chars
inlen the length of in
- Returns
- 0 if success, -2 if the transcoding fails, or -1 otherwise The value of inlen after return is the number of octets consumed as the return value is positive, else unpredictable. The value of outlen after return is the number of octets consumed.